Normal view

How AI helps scientists design the next generation of medicines

Designing and developing a new medicine is an expensive, failure-prone scientific challenge. A new drug can take many years to develop, at the cost of a significant investment. And even then, most possible candidates never reach the patient. For biologic medicines, therapies made from engineered proteins rather than synthetic chemistry (which are often used to treat conditions across most major acute and chronic diseases), the complexity is even greater.

Scientists explore vast quantities of possible molecules, looking for the rare few that will bind to the right target, remain stable in the human body, and be manufacturable at scale. Today, AI is speeding up these processes and has quickly become a core part of the infrastructure in pharmaceutical R&D.

AI-assisted design is a growing part of how biologic drug candidates are developed, and companies like AstraZeneca are actively building its engineering teams to push this further. “Everything we do, whether it’s design, make, test, or analyze, is now computationally enhanced,” says Puja Sapra, senior vice president and head of R&D biologics engineering and oncology targeted discovery at AstraZeneca. “The cycle times are getting shorter while productivity and innovation increase.”

Sapra explains that AstraZeneca’s approach follows a build-measure-learn loop. AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed. Scientists then focus lab resources only on the top-ranked candidates. This leads to a tighter feedback cycle with fewer dead ends, faster iteration, and the ability to go after disease targets that were previously considered untreatable by medicine. Because the number of possible molecular combinations far exceeds what any human team can systematically explore, using AI to narrow and refine the options for testing has become a major focus in biologics drug design.

Navigating complex drug design problems

Beyond accelerating timelines, AI is also being applied to the discovery of entirely new classes of medicines. Traditional biologics typically target one disease pathway. The next generation of drugs can hit multiple targets simultaneously or precisely deliver therapeutic payloads to specific cells. Achieving this requires optimization across many variables at once. Looking ahead AI-driven models could help design these increasingly complex, multi-specific biologics, explains Puja Sapra. “For example,” she continues, “such models could help identify which two or three targets to prioritize based on the underlying biology, then optimize across multiple parameters to balance a molecule’s potency, stability, manufacturability, and safety.” “Drugging the undruggable is becoming a reality,” Sapra says. “These technologies will eventually enable us to develop medicines against targets once thought impossible to reach. The potential for benefit to patients is remarkable.”

The data moat

McKinsey estimates that generative AI, combined with other computational tools, could cut drug discovery timelines by as much as 50%. But every AI model is only as good as its training data. In drug discovery, that means ample quantities of high-quality biological data. Experiments can provide a rich source of such data. Whether they succeed or fail, each experiment generates a signal about what does and does not work.

“Data is our differentiator,” says Sapra, explaining how the company’s datasets are proprietary and multimodal and include molecular structures, binding measurements, safety profiles, and manufacturing outcomes. “We’ve built an intentionally diverse portfolio across multiple disease areas and drug types. All of that data empowers us to fine-tune frontier AI models with richer, more representative training sets.” She continues, “Further, we have invested in deep screening technologies to generate additional datasets required in volume to constantly refine and validate our models.”

Building an autonomous discovery engine

To bring all of that data together in one place, AstraZeneca is building what it calls a “lab of the future” facility in Kendall Square, Cambridge, Massachusetts where AI and robotic automation will be able to form a continuous, closed-loop discovery system. “Where a self-driving car uses sensors and models to navigate its environment, this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data,” explains Sapra. That data feeds directly back into the models, accelerating each subsequent cycle.

“Throughout, scientists will remain central to the process, providing the oversight, judgement, and strategic direction that ensure outputs are explainable, tolerable, and directed toward potential patient benefit,” she adds.

Eventually, automated high-throughput systems will be able to make and evaluate thousands of molecular interactions on a weekly basis. “This will generate AI-ready data at a scale that traditional workflows cannot match,” Sapra says. “Robotic sample handling, automated quality checks, and integrated data pipelines also have the potential to help accelerate early drug development timelines significantly.”

The next frontier: Generating medicines from scratch

Ultimately, Sapra says, the end-state vision for AI in biologic drug discovery is what the field calls “de novo” design. For this, the goal is for AI to generate entirely new protein sequences that precisely fit the desired drug properties. This includes designing the structure, predicting safety, how it will behave in the body and how to make it manufacturable.

“The field is making great progress toward a completely AI-generated biologic, designed from scratch all the way to a clinical candidate,” Sapra says. “As we continue to leverage frontier models and fine-tune them with the right datasets, we bring ourselves closer to this reality. I believe it will come. It’s a matter of time.”

Several key elements are needed to reach this point, however. First is richer and more standardized training data across the industry. Second, robust evaluation benchmarks for AI-generated candidates. And third, teams that know how to work at the intersection of machine learning and biology. Of all the prerequisites, however, safety prediction may be the most consequential, and perhaps the least discussed, Sapra says.

“One of the hardest problems in de novo design is predicting whether a computationally generated molecule will be safe in the human body,” Sapra explains. AstraZeneca is tackling this with what amounts to virtual clinical trials. These are advanced cell systems and micro-scale organ models that function as physical testbeds, paired with AI that learns from their outputs.

 “These systems have the potential to generate enhanced biological signals without traditional testing bottlenecks, and they’re a critical missing piece in closing the loop between AI-generated designs and clinical-ready candidates,” Sapra adds.

A shift currently underway is the move toward agentic AI systems that can simultaneously generate molecule candidates and predict how efficacious and safe they are likely to be. These autonomous workflows can connect disease-level insights directly to molecule design, bridging what were previously separate data silos. “The complexity of the biology goes hand-in-hand with the design of the molecule,” summarizes Sapra.

Human talent unlocks AI potential

The transformation underway in biologics is not just about technology. “With more autonomous systems, human oversight remains at the heart of this approach—ensuring explainable and ethical AI for the benefit of patients,” says Sapra.

For scientists, working with AI is a collaborative process. “Scientists will work hand-in-hand with these model systems,” she says. “There will be a world where models will design molecules, then scientists will work with the systems to test those molecules and put all that data together.” Through this process of human checks, balances, and judgement calls, the models will evolve and constantly improve, ultimately with potential to benefit patients.

For engineers, designing and building effective systems ready for human-AI collaboration will mean ensuring high levels of model transparency and explainability. According to Sapra, AstraZeneca’s engineering teams include data scientists, automation specialists, and AI engineers, who are developing systems that act as “thinking partners” rather than black boxes. “Engineers are designing systems that generate, validate, and learn at speed. And the problems are genuinely hard: Multimodal data fusion, closed-loop optimization, uncertainty quantification, and interpretability at the point of clinical decision-making,” she adds.

In taking on such technically demanding challenges, engineers and scientists have the opportunity to contribute to the research and development of potentially life-changing treatments for many diseases, says Sapra. “The biologic medicines we can develop today, and those we’ll design tomorrow, depend on combining world-class AI and engineering talent with deep scientific expertise.”

This article has been initiated and funded by AstraZeneca.  Z4-85058, July 2026.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Advancing next-gen AI with materials science innovation

The conversation about AI often centers on algorithms, computing power, or huge investments in new semiconductor fabrication plants and hyperscale data centers. But beneath each of these advances is another layer of innovation that makes them possible: advanced materials.

Every new generation of AI technology demands more processing power, more memory, greater energy efficiency, and higher reliability. Every increase in computing performance increases the physical demands placed on the systems that make and run AI.

Delivering these gains depends not only on advances in chip design and system architecture, but on advances in the materials that enable them to perform under extreme conditions.

As AI continues to push the physical limits of semiconductors and data center infrastructure, advanced materials are no longer simply supporting innovation in this area; they are defining the limits of what is possible.

Performance first

Advanced materials exist to solve performance challenges. As AI raises the bar, these challenges are becoming more demanding.

Manufacturing a semiconductor chip today requires thousands of tightly controlled process steps, with almost no room for error. Tiny variations in temperature or chemical instability can create defects that reduce yield and drive up manufacturing costs. With every new generation of semiconductor chips, manufacturers seek advanced materials that can deliver greater purity, higher chemical and plasma resistance, and better stability under increasingly harsh operating conditions.

These are familiar engineering challenges being pushed to new extremes. And it’s here that materials innovation makes the difference with continuous advances in polymers, elastomers, specialty fluids, and other advanced materials that make each new generation of technology possible.

For materials companies, it’s not about reinventing semiconductor manufacturing but about ensuring the materials supporting the industry continue to evolve alongside it. This same principle applies beyond the semiconductor fabrication floor. As AI workloads become more demanding, the physical infrastructure that powers them is evolving rapidly.

Increasing computing density is transforming data center design, driving the need for more sophisticated thermal management, higher-voltage power architectures, increased data storage, and faster, more reliable data transmission. Every part of the system is under greater pressure, from cooling and power management to critical electronic components, such as connectors, capacitors, and hard disk drives.

At Syensqo, we’re building on our expertise in electronic and electrical components, along with insights from other markets, to meet these emerging needs.

For example, as data centers shift to higher-voltage architectures and greater power density, many of the materials challenges we face closely mirror those of electric vehicles. Fluid-circulation know-how from semiconductor and automotive coolant systems, for instance, can be adapted to direct liquid-cooling designs for AI servers. By transferring knowledge across markets, we can accelerate new power and thermal management solutions while supporting the reliability required by next-generation AI infrastructure.

Whether we’re talking about semiconductor fabrication or hyperscale server farms, the challenge for materials science companies is the same: enabling greater performance without compromising reliability.

A new definition of what performance means

While performance remains the first priority, the way performance is defined is changing.

In addition to meeting the increasingly demanding technical requirements of next-generation semiconductors and data centers, there is now an expectation that these materials are developed and manufactured more responsibly.

Perfluoroelastomers, for example, are used to seal semiconductor manufacturing equipment. These materials operate under extreme temperatures, aggressive plasma, and highly reactive chemicals.

To make the process more sustainable, at Syensqo, our next generation of perfluoroelastomers use a fluorosurfactant-free manufacturing process. Our goal was to make a better-performing material, produced in a better way, ensuring manufacturers no longer have to choose between higher performance and a more responsible way of producing the materials that enable it.

This approach reflects a broader reality across the industry.

New materials aren’t adopted simply because they are new. Qualification can take years, and manufacturers only make changes when a material solves a genuine engineering challenge or enables new technology.

Performance remains the price of entry. The difference today is that the definition of performance has expanded. Success increasingly depends on delivering technical excellence through more responsible manufacturing from the outset.

Accelerating the pace of discovery

As the performance bar rises, the way we innovate must evolve with it.

Developing advanced materials has traditionally involved a lengthy process of hypothesis, synthesis, testing, and iteration. While this process remains unchanged, new digital tools are helping researchers move through these cycles faster. By helping researchers identify the most promising candidates earlier, AI can reduce the number of physical experiments required and accelerate the earliest stages of materials discovery.

AI isn’t replacing scientific expertise. It’s helping scientists apply that expertise more effectively, allowing them to spend less time searching for answers and more time solving the industry’s toughest challenges.

At Syensqo, we’re putting this approach into practice through use of several AI tools, including the Microsoft Discovery platform, which are helping researchers identify and evaluate promising molecular candidates for next-generation heat transfer fluids, used in semiconductor manufacturing and data centers.

AI helps our researchers rapidly identify and evaluate promising molecular candidates based on the properties they need to achieve. This allows us to focus laboratory work where it has the greatest potential to deliver results, accelerating discovery and reducing the time needed to turn promising materials into solutions customers can qualify and deploy.

The journey from laboratory discovery to a qualified material will always require scientific expertise, rigorous testing, and close collaboration with customers. But by accelerating the earliest stages of discovery, AI can help materials innovation keep pace with the evolving needs of industries such as semiconductors, electronics, and data centers.

Progress is earned

The future of artificial intelligence will depend on better algorithms, more powerful chips, and larger computing infrastructure. But sustaining that progress will also require advances in the materials that make those technologies possible.

Whether in semiconductor manufacturing or AI infrastructure, progress is earned. Every new generation of technologies raises the bar, and every new material must prove it can deliver the performance, reliability, and efficiency needed before it earns its place.

For materials companies, that remains both the challenge and the opportunity.

This content was produced by Syensqo. It was not written by MIT Technology Review’s editorial staff.

Arm and Google offer a smarter option to run agentic AI workloads

Warp speed light streaks radiating outward on blue background

As enterprise leaders start deploying agentic workflows, they must establish the infrastructure to build and run them, one capable of fluidly routing a diverse set of workloads across the most efficient compute resources.

This requires the ability to manage heterogeneous infrastructure, utilizing high-performance accelerators for large-scale training and inference, and utilizing CPUs for the critical orchestration layer of agentic AI. As autonomous agents become more prevalent, CPUs are ideally suited for managing agent state, semantic routing, tool selection, and spinning up secure, isolated sandboxes to safely execute untrusted generated code.

The Google Axion advantage

Google Cloud, with its workload-optimized Compute Engine portfolio, which includes general-purpose and specialized offerings, shines in addressing this need.

Google Axion processors within this portfolio comprise a family of custom Arm processors engineered for performance, efficiency, and versatility, with a feature set that supports general-purpose workloads, CPU-based AI workloads, and other specialized tasks requiring Arm-native compatibility and direct hardware access.

Axion is Google’s first custom Arm-based server CPU, introduced in April 2024. It is designed specifically for hyperscale cloud and AI-era data center workloads. 

Axion also leverages more than a decade of Google’s custom silicon innovation. This enables Google to more readily incorporate customer feedback into chip designs and address the more general, though complex, needs of CPUs. 

Matching workload type to the processor

Bhumik Patel, Director of Software Ecosystem Development at Arm, says the key to all of this is to match the workload type as closely as possible to computing capacity. CPU-powered cloud instances are a practical option for certain AI workloads, particularly those with smaller datasets or less complex models. 

“Agentic tasks such as orchestrating, talking to APIs, and memory management are all ones CPUs are good at, so it’s a distributed and concurrent AI workload,” Patel tells The New Stack. Intelligent workload-processing apportionment makes agentic AI more cost-effective and efficient than running all workloads on a single compute type.

This efficiency is quantifiable. The Google Kubernetes Engine Agent Sandbox running on Google Axion N4A provides up to 30% better price performance than the next hyperscale cloud provider, says Google’s Mo Farhat, Axion Group Product Manager. The GKE Sandbox is an open-source Kubernetes-native primitive designed to execute untrusted AI-generated code safely. 

“Agentic tasks such as orchestrating, talking to APIs, and memory management are all ones CPUs are good at, so it’s a distributed and concurrent AI workload.”

Intelligent workload decoupling makes agentic AI significantly more cost-effective. Google Cloud’s fluid computing foundation enables engineering teams to reserve specialized accelerators strictly for heavy reasoning and generative workloads, while leveraging Axion CPUs for high-concurrency orchestration and context management.

Secure execution with the GKE Agent Sandbox

As agents begin to generate and execute dynamic code autonomously, security is non-negotiable. Running AI-generated code directly in a standard cluster poses severe security risks, as untrusted code could potentially access other apps or the underlying cluster node.

The Google Kubernetes Engine (GKE) Agent Sandbox resolves this by providing an isolated environment for safely executing untrusted code. Running on Axion-powered N4A instances, the sandbox provides up to 30% better price performance than comparable workloads on other hyperscalers.

The vertical stack isolates sensitive tasks at the kernel level with sub-second latency.

The vertical stack isolates sensitive tasks at the kernel level with sub-second latency.  GKE Agent Sandbox natively supports gVisor (an open-source application kernel developed by Google that acts as a secure sandbox for containers) and default-deny Kubernetes network policy. Agent Sandbox provides pluggable interfaces for open-source sandboxes, such as Kata Containers, enabling users to customize their kernel isolation. 

Powered by gVisor technologies with software support from Arm’s architecture, the sandboxes intercept and validate system calls before they reach the host kernel. These isolated execution environments enable deployment of autonomous systems at scale without sacrificing performance or operational agility.

To manage resources efficiently when agents sit idle, GKE Pod snapshots allow users to save and restore the exact process state of sandboxed environments. This functionality provides four major architectural benefits:

  • Fast startup: Reduces sandbox startup time by restoring from a pre-warmed snapshot rather than initializing from scratch.
  • Long-running agents: Pauses sandboxes that take a long time to run and resumes them later—or moves them across nodes—without losing progress.
  • Stateful workloads: Persist an agent’s context, such as conversation history or intermediate calculations.
  • Reproducibility: Captures a specific state to use as a baseline for spinning up multiple new sandboxes.

Getting started

As token generation, autonomous workflows, and continuous agent interactions grow exponentially, relying exclusively on accelerator-backed stacks for every task will become financially and architecturally unsustainable.

The combination of CPU and accelerator execution accounts for bursts in agent activity and unpredictable demand spikes by eliminating the inference tax. Google Cloud’s full-stack advantage enables organizations to deploy the right machine for the job. 

By using Google Axion and GKE Agent Sandbox, builders can optimize total cost of ownership and security while maintaining the performance required for AI agents.

Learn more about Google Axion.

The post Arm and Google offer a smarter option to run agentic AI workloads appeared first on The New Stack.

Building a Foundation Stack for General-Purpose Robots

13 July 2026 at 10:19


This article is brought to you by X Square Robot.

Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the field does not yet agree on what it is.

X Square Robot, a Chinese embodied-AI company, has made an unusually explicit bet. It argues that the recipe is an integrated stack, spanning the data a robot learns from, a world model for predicting changes in the physical world, and an action model that brings together perception, planning, reasoning, and decision-making to generate executable robot behavior. The company also believes that the stack should be built and released in the open.

X Square Robot shares its vision of bringing robots into real homes.X Square Robot

X Square Robot’s embodied AI stack

What holds the stack together is a small set of principles rather than a single overarching model.

  • The first is that the basic unit of robot data is an interaction, not a trajectory; a demonstration is successful only if it changes the world as intended, not simply because the joints moved.
  • The second is that pretraining should yield usable capability, not just an initialization for later fine-tuning.
  • The third is that behavior should be modeled around physical events rather than fixed slices of time.

These principles make the layers interdependent, since the same robot-free data that trains the action model is also structured to feed the world model. It is worth being precise, though. The company describes the world model and the action model as complementary but independent model families that share a code base. Both sit within its broader World Unified Model, which it has presented as an architecture for training vision, language, action, and physical prediction together.

Robot learning data: Engineering for quality and cost, not scale

For the X Square Robot team, one of the biggest constraints on general-purpose robots is the cost and quality of interaction data, not the number of parameters. To address that, the company built its Universal Manipulation Interface (UMI) data collection system, QUANXTA Zero Series. It works by collecting demonstrations from people wearing a rig with dual grippers rather than teleoperating a robot. This approach is not itself new, and builds on established methods for robot-free data capture. What sets it apart are two engineering choices.

Person using VR headset and handheld controllers to teleoperate a dishwashing robot system X Square Robot emphasizes data quality control, recording trajectories and replaying them on a real robot, with only those that actually complete the task counted as valid.X Square Robot

The first is quality control, and it is the most distinctive part. Rather than accepting recorded trajectories as they are, the system runs a closed inspection loop, and its notable step is physical playback. A sample of trajectories is replayed on the real robot, and only those that actually complete the task count as valid. That makes the validity rate a measured quantity rather than an assumption. For example, a gripper that closes a fraction of a second too early still looks like a grasp in the data, yet it has pushed the object away, so it shouldn’t be classified as valid. A smaller clean dataset can be worth more than a larger noisy one.

The second choice is how lower-cost human data and scarce robot data are combined. The company pretrains on a large volume of robot-free demonstrations to build general representations, then adds a small amount of real-robot data as an anchor to the specific machine’s dynamics. It reports that this reaches performance comparable to an all-robot dataset at roughly a 20-fold lower cost of collection, driven mainly by how much cheaper the wearable rig is than a teleoperation setup.

The resulting dataset is deliberately model-agnostic, formatted to feed both action models and world models. The caveat is that the strongest results are measured on the company’s own robots and data-collection pipelines. Broader independent testing will help confirm and extend these promising results across a wider range of settings.

A world model organized around events

In developing its world model, called WALL-WM, X Square Robot took a differentiated approach. Most action models predict a fixed-length chunk of motion from the current image and instruction. That is convenient, but it segments behavior into fixed-duration windows, so the boundaries fall where elapsed time dictates rather than where one action ends and the next begins. WALL-WM instead treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.

Collage of robot arms manipulating kitchen objects with charts of multimodal AI performance X Square Robot’s world model, called WALL-WM, treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.X Square Robot

WALL-WM’s design reflects a specific concern about not discarding what large video models already know. To achieve that, a text-to-video model is coupled to a freshly initialized action network that reads from the video features without overwriting them, which preserves the visual prior. From that one process, it offers two modes. An event mode runs in variable-length segments and suits reasoning over long horizons, while a fixed-length mode produces the steady, real-time output a controller needs. That places WALL-WM between mainstream chunk-based action models and pure video world models, keeping the predictive character of a world model while still yielding executable control.

In a series of experiments, the company relied on a generalization test that is more specific than most. A model trained on a limited dataset was evaluated on long-horizon tasks in unseen settings and, on the company’s real-robot benchmark, reportedly outscored baselines that had been fine-tuned on related data. That is a meaningful result if it holds. For now, it is measured on the company’s own benchmark. With the code now being released, the broader community will have the opportunity to test, reproduce, and build on them across more settings.

A policy that runs before fine-tuning, and action tokens with meaning

The action layer carries two connected ideas. The first is a requirement the company sets for itself with Wall-OSS-0.5, its vision-language-action model: The pretrained model should run on a real robot before any task-specific fine-tuning.

The interest is less in the scores than in the design behind them. The model trains three objectives together, namely discrete action tokens, language grounding, and continuous action generation. And it keeps gradients flowing through all of them rather than freezing parts of the network as some rival designs do. It’s also a more strict method, since it reports untuned behavior such as approaching, grasping, and recovering, including on a deformable task held out of training.

Dashboard of robot training metrics with charts and photos of a robot sorting objects As part of X Square Robot’s Wall-OSS-0.5 vision-language-action model design, the pretrained model should run on a real robot before any task-specific fine-tuning. X Square Robot

The second idea is the action interface itself, called X-Tokenizer. Most systems that turn continuous motion into discrete tokens produce codes that the language model cannot interpret. X-Tokenizer reframes tokenization as learning a semantic interface, so that the top-level code stands for the intent of a motion while lower-level codes carry finer detail, all aligned with the language model’s own features.

A useful consequence is stability. Adding noise to an action barely moves the intent code, which is what lets one tokenizer to be reused across robots without re-tuning. The tokenizer inside the production action model is a related variant of this approach. Together, the two ideas give the action layer something rather powerful: capability that transfers.

The future of embodied AI stacks

X Square Robot is betting that its unique approach combining three layers, each specialized in solving a key part of the problem, will stand out from other embodied AI stacks. The physical-playback step that grounds data quality is uncommon and sensible. The reframing of world modeling around events, with one backbone serving both reasoning and control, is a genuinely distinct approach. And the pairing of a deployable pretraining standard with a tokenizer designed as a semantic interface gives the action layer unusual coherence.

X Square Robot’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.

The next phase will bring broader validation. Much of the current evidence comes from X Square’s own robots and benchmarks. With the world model code now being made public, and as the community begins to test, reproduce, and build on the work, the reported capabilities will be tested across more robots, tasks, and settings.

X Square Robot’s recent funding rounds reflect similar confidence. The company’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.

What’s next for X Square Robot

To learn more about its future plans, the following Q&A with the X Square Robot team further explores the company’s technology, strategy, and vision.

What made now the right moment, technically, to commit to this stack? What recently became possible that wasn’t possible a couple of years ago?

It is not one breakthrough but several trends maturing together. Foundation models gave us a shared representation across vision, language, and action, so we can model what a robot sees, what it is asked to do, and how its actions change the world in one framework, rather than as separate perception, planning, and control modules.

Compute and infrastructure are finally sufficient for large-scale pretraining over long-horizon, multi-embodiment data. Just as importantly, we realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical. The useful question is no longer how to predict a few seconds of video, but how to understand the ways actions change objects, contacts, and task states. Two years ago these ingredients existed separately. Today they are mature enough to work as one system.

“We realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical.”

Your data system captures demonstrations with a wearable VR rig and custom grippers rather than teleoperating robots. What was wrong with standard teleoperation?

Teleoperation is built around controlling the robot. It forces the operator to work within the machine’s kinematics, latency, and viewpoint, and the resulting demonstrations are slower, stiffer, and less diverse. We built our system around capturing human skill instead. Manipulation is really about contact, timing, finger coordination, and recovery, not just the path the hand takes, and a wearable rig records those before the behavior is compressed onto one particular robot. It also breaks teleoperation’s expensive scaling law, in which every demonstration needs a robot.

People can generate rich data independently of any robot, and the crucial property is that those demonstrations can still be replayed and executed on a physical robot through the model. Mobility is convenient, but that replay is the real point, because it is what lets the same data be reused across different platforms.

Robot and person loading a washing machine together in a modern laundry room. In X Square Robot’s approach, demonstrations can be replayed and executed on a physical robot through the AI model, allowing the same data to be reused across different platforms.X Square Robot

X Square Robot reports that its pipeline has roughly an 85 percent data-validity rate. Why is quality control such an underrated bottleneck?

Because errors in robot data are far more expensive than in language data. A small timing or contact error can change what a demonstration means. If a gripper closes a fraction of a second too early, the motion still looks like a grasp, but physically it has pushed the object away. A dataset that mixes failures and accidental successes teaches ambiguity, not skill, because the real unit is the interaction, not the trajectory.

So we run automated inspection, kinematic checks, and physical replay, where we play a sample of trajectories back on the real robot and count only the ones that actually complete the task. Data quality sets the ceiling on how good a policy can be. In our experience a smaller, cleaner dataset often beats a much larger, noisier one, which is why we treat quality control as part of the model, not a preprocessing afterthought.

The model runs in both “event mode” and “chunk mode.” When does each matter?

Both matter, for different reasons. The physical world changes through events—when contact occurs, a grasp forms, or an object slips—not in fixed-frame windows. Event mode concentrates the model’s attention on those moments, and it matters most for long-horizon tasks, like clearing a table, where progress is a sequence of semantic events rather than a smooth stream. It runs in variable-length segments that follow the task rather than a clock. Chunk mode matters for deployment. Real controllers need a stable, real-time interface, and fixed-length chunks integrate cleanly with existing control systems.

We organize learning around events in the first place because a fixed window can split one motion in half or merge two together, which turns training into short-horizon pattern matching and weakens the model on long tasks. So the world model’s job is to connect event-level understanding, which is where the reasoning happens, with a fixed-length output a real robot can actually run.

Why make “deployable before fine-tuning” the criterion?

Pretraining should produce capability, not just a good starting point. If a model is only useful after heavy fine-tuning, then most of the intelligence still lives in the downstream supervision, not in the foundation model. Deployable before fine-tuning is a more honest test of what pretraining actually learned. A well-pretrained robot should already know how to approach, grasp, move, avoid obstacles, and correct itself. Fine-tuning should adapt it to a specific task or robot, not create the ability from nothing. It is also a practical requirement. A robot in a home or a workplace shouldn’t need a brand-new dataset and a new policy every time the task changes, so a foundation model that already carries general skill, and some ability to recover, is the minimum bar for something genuinely useful in the real world.

What is the most challenging part of cross-embodiment learning?

Robots differ in control frequency, delay, compliance, sensing precision, and contact dynamics, so the same instruction can require different action decompositions and recovery strategies, and a behavior that works on one arm cannot simply be copied to another. Cross-embodiment learning needs an intermediate abstraction, lower than language but higher than joint angles: how you approach an object, how you make contact, how you apply force, and how you recover from a mistake.

When we say cross-embodiment, the main capability we mean is multi-embodiment generalization: transferring across robots, training on many embodiments at once, and adapting to different kinematics. Human-to-robot transfer and other techniques are specific approaches to that goal.

“A robot in a home or workplace shouldn’t need a new dataset and policy every time the task changes. A useful foundation model should already carry general skills and the ability to recover.”

What would you most like to see other researchers attempt to reproduce or stress-test?

Three things, above all. Whether event-level representations really generalize beyond our own datasets, across more tasks, scenes, objects, embodiments, and failure conditions. Whether pretraining stays effective on robots the model never saw during training, or whether its capability is still too tightly coupled to what it has already seen. And whether real-robot evaluation can become a shared language for the field, so that we compare not just success rates but the reasons systems fail, where an instruction was misread, where perception broke down, or where recovery fell short. Robotics has been driven too often by impressive demonstrations, and real progress comes from results that are reproducible and diagnosable.

What capability is still missing before robots become dependable in homes?

Benchmarks measure competence, like whether a model can finish a task. Homes demand reliability, safe and consistent operation over time in a place that changes every day, with objects moving, instructions that are vague, and people interrupting. The missing piece is not a higher one-time success rate: it is robust recovery. A dependable home robot has to know when it is uncertain, when to slow down, when to ask for help, and how to bring the world back to a safe state after it drops something or misunderstands a request.

In a real home, failure recovery matters more than raw success, because the home does not reset itself. Homes also demand careful personalization, learning a household’s routines and preferences over time, with safety and trust as first principles. That combination, not any single skill, separates a capable demonstration from a robot people can live with.

Humanoid service robot stands by a table in a modern living room. X Square Robot’s approach is that, in a real home, failure recovery matters more than raw success, because the home does not reset itself and it demands careful personalization, with safety and trust as first principles. X Square Robot

How do the open-source components fit into X Square Robot’s World Unified Model direction?

We see these releases as layers of the World Unified Model direction rather than isolated projects. Wall-OSS-0.5, the action model, asks whether an open vision-language-action model can gain directly measurable capability from large-scale pretraining, so it is the capability layer. WALL-WM, the world model, asks how a robot should understand change in the world, shifting from fixed windows to event-level modeling, so it is the representation layer. The data system supplies the interaction data that both of them learn from.

Together they form a loop in which models produce capability, world models organize understanding, and the open-source community drives reproduction and improvement. World Unified Model is the broader architecture those layers support, bringing vision, language, action, and physical prediction together.

We are releasing these pieces openly because embodied intelligence cannot be solved by one organization; it needs many embodiments, many real tasks, and broad feedback, and the long-term goal is a stack that keeps learning and ultimately moves robots from laboratory demonstrations toward reliable everyday use.

The foundational elements of AI architecture that IT leaders need to scale

With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces risk, leaving IT leaders to wonder which investments will prove valuable even six months into the future.

Returning to the foundational elements of AI architecture—the structural framework required for deploying and managing reliable, integrated AI systems at scale—allows technology leaders to make astute decisions today while supporting a future of AI agents that can retrieve information, make decisions, and execute complex workflows across systems.

Four elements of AI architecture you can count on

The following capabilities provide a stable compass on the path to production-ready deployment, regardless of how the underlying technology evolves.

1. Prepare data for AI at scale

Models are only as reliable as the data they can access, and poor data quality leads to AI hallucinations, bias, and unreliable outputs.

Most enterprises rely on legacy systems, inconsistent data structures, fragmented ownership, and incomplete datasets, making it difficult to scale AI effectively. Powerful as it is, AI itself cannot solve these underlying data problems.

As Adnan Adil, CIO of Elastic, explains: “The data is a durable part of AI architecture because without it, these models won’t run, won’t provide the right context, or won’t give the right level of services that we’re looking to implement.” Industry surveys consistently cite data quality as one of the greatest barriers to AI success. “The data quality has to be good; otherwise, the user loses confidence in the system,” says Adil.

An effective AI strategy begins with connecting data across the organization and ensuring it is organized, accurate, governed, and accessible in real time. These considerations are most effective when built into models and architecture from the start. Scalable data architecture allows AI systems to evolve alongside the business and connect reliably to the internal information needed to deliver meaningful value.

Gartner predicts that companies will abandon 60% of all AI projects through 2026 if they are not supported by AI-ready data. Avoiding that outcome includes clear data standards and ownership, clean and labeled data, and pipelines that support real-time retrieval.

2. Use context engineering to deliver the right data to every AI query

Context engineering ensures that the model draws on the most pertinent information for each query, selecting and organizing the data needed to produce accurate answers efficiently.

Effective context engineering shapes the inputs that guide AI reasoning and action. While prompt engineering focuses on how a request is worded, context engineering designs the entire information environment around the model: retrieving the right data and presenting it in a structured, machine-readable way. Many organizations are discovering that reliable AI depends as much on context quality as on the strength of the model.

Context engineering relies on a modernized, unified data foundation as well as retrieval and memory systems such as retrieval augmented generation (RAG) and vector databases. It also requires careful prioritization to determine what information matters most, what should be excluded, and when different types of information should be used. Feeding models too much context can dilute relevant details, increase costs, and slow response times.

“Minimum context, correct and current data, and machine-readable information are critical to effective context engineering,” Adil says.

3. Build AI governance and LLM observability in from the start

Strong governance and LLM observability help organizations maintain control over how AI systems use data, monitor system performance, and identify problems before they affect operations.

In the absence of clear controls around retrieval, workflows, and model usage, AI systems often process far more information than necessary. This inefficiency also drives up operating costs by requiring additional computing resources, often reflected in higher token consumption and API charges.

Governance also works in tandem with robust security. AI expands the attack surface, introducing risks such as prompt-based data leakage, model vulnerabilities, and adversarial inputs. Protecting sensitive information requires strong access controls, monitoring, and oversight.

Adil notes that essential controls — including those related to security, granular cost management, project controls, data security, and architecture—are frequently insufficient.

For governance systems to support transparent, compliant, trustworthy, and cost-effective AI, organizations cannot leave them as a layer to add later. Governance structures need to be embedded into architecture, workflows, and decision-making processes from the outset.

When governance is established from the start, it enables robust observability. Observability helps organizations understand how AI applications are performing in practice. Mechanisms for LLM observability and benchmarking allow teams to assess accuracy and utility over time, monitor adoption patterns, and adjust systems as conditions change. Observability also helps organizations gain trust by increasing visibility of model performance, behavior, and failure points.

Furthermore, observability is essential to get ROI of AI initiatives, as the benefits of it are often indirect and business value depends heavily on how systems are adopted and used. Real-time visibility into AI behavior allows organizations to measure performance against expectations, identify gaps between intent and reality, and continuously refine systems as requirements evolve.

In a 2026 report from Elastic, 85% of IT decision makers expect to enable LLM observability for their internal generative AI apps.

“Observability is actually huge. We can use observability data for cost control, decision-making, and engineering efficiency,” Adil says.

4. Keep humans in the loop

The thoughtful design, integration, and governance that maximize AI value demand specialized in-house expertise. Nearly 70% of respondents in Deloitte’s 2025 Tech Executive Survey report plan to grow teams in direct response to generative AI, a clear contrast to widely reported AI-related cuts. Adil agrees: “We think the people aspect is largely what’s going to make AI impactful going forward.”

As AI systems become more embedded in operations, organizations need people who can govern workflows, evaluate outputs, redesign processes, and adapt systems as conditions change. Evolution toward increasingly autonomous tools requires teams skilled in prompt engineering, orchestration, and change management. 

Talent adept at critical thinking and prepared to adapt with technology’s rapid advances will be in high demand. Although turnover brings in fresh thinking, it also presents high costs in system continuity, institutional understanding, and innovation. Human-centered strategy needs to be built into AI execution stages to ensure smooth implementation. 

As Adil says, “Many aspects of the stack are moving very, very fast, but institutional knowledge and the ability to adapt remain durable.

Thoughtful AI investment for future growth

As AI systems evolve from single-task assistants to increasingly autonomous agents, the organizations best positioned to benefit will be those that invest in the underlying systems, governance, and expertise that make AI reliable at scale.

Tech leaders who focus on these fundamentals can move effectively from experimentation to reliable, production-level deployment in the medium term, confident that these elements will remain relevant and adaptable amid constant advancements.

“We fundamentally believe that with these tools, velocity of work will get much faster,” Adil says. “We are really focused on how we can do work with these tools in ways we had not thought of before.”

Learn more about how Elastic is building an AI-first enterprise with these core foundational components.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Achieving operational excellence with AI

Frameworks like Lean Six Sigma and business process management (BPM) first gained traction because they promised clarity in the chaos—a structured way to bring order to messy, sprawling operations. Lean Six Sigma emphasized statistical rigor and quality control; BPM created end-to-end maps of how work should flow across departments. Both offered a repeatable way to embed habits of measurement, analysis, and accountability into day-to-day company culture.

But today, those time-tested playbooks are evolving as companies seek to embed AI into established process excellence methodologies. By some estimates, the market for AI-powered process optimization is projected to exceed $113 billion within the next decade. In one study, a full 88% of business leaders anticipated increasing investments into AI-infused process intelligence in the next 12 to 18 months.

Yet without the right foundations, many of those investments may not fully deliver on their potential. Companies that already operate with discipline have an edge. They can channel new tools into proven systems rather than bolting them onto shaky foundations. Organizations with mature process disciplines are also better positioned to translate AI ambition into real outcomes, as they are already accustomed to data-driven decision-making and process discipline—precisely the cultural foundation AI systems need to deliver value.

Simply put: AI can accelerate process excellence, but existing process excellence is what makes AI truly impactful. Technology and process are no longer separate levers, and only organizations that pull them together stand to realize the full value of both.

Download the full report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Building the foundation for an autonomous enterprise

Artificial intelligence may have captured the public imagination through chatbots and image generators, but some of its most consequential use cases are unfolding far from consumer-facing tools. In industries where physical infrastructure, operational continuity, and safety are paramount, AI is becoming a core operating layer. With its sprawling industrial systems and constant stream of operational data, the energy sector offers a glimpse into what that future could look like.

At Woodside Energy, AI adoption did not begin with generative models or enterprise copilots. The company has spent years building predictive analytics, optimization systems, and machine learning tools across exploration, drilling, maintenance, and plant operations. “We’ve always had very large volumes of operational data coming from the equipment and the plants and the assets that we operate,” says the company’s vice president for digital Andrew Melouney. “Those have created really clear, quite high-value use cases for us.”

That long-term investment in infrastructure and governance is now enabling a broader shift toward agentic AI systems that can support complex industrial workflows. Rather than replace human operators, Woodside designs AI systems to augment expertise in high-stakes environments. A prime example is its “Startup Advisor,” an AI copilot that helps operators manage the complex process of starting liquefied natural gas (LNG) plants. “We’re really thinking about, how does it support the people in the organization in terms of empowering them to make better decisions, to make faster decisions,” Melouney explains.

The company’s approach reflects a wider evolution taking place across industrial AI: graduating from isolated experiments to enterprise-wide systems built on standardized platforms, governed data, and repeatable deployment patterns. That transition, Melouney argues, requires organizations to rethink both their technology stacks and how work itself gets done. “We’re not just bolting AI onto an existing process,” he says. “We’re deeply thinking about how that work needs to be reimagined.”

Melouney’s motto has become: “Think big, prototype small, and scale fast.”

As AI systems become more autonomous and interconnected, the companies poised to succeed may be those that spent years building the operational foundations beneath the hype.

“Our ambition is really for an autonomous enterprise, where we have agents with agency that are able to really deeply interact with our core workflows,” says Melouney.

This episode of Business Lab is produced in partnership with Infosys.

Full Transcript:

Megan Tatum: From MIT Technology Review, I’m Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace.

This episode is produced in partnership with Infosys.

Now, when people think about artificial intelligence, they often picture chatbots or productivity tools, but some of the most sophisticated and high impact uses of AI are actually happening far from consumer apps, inside complex industrial environments where safety, reliability, and physical systems matter. The global energy sector is a prime example.

Companies like Woodside Energy, a global energy producer headquartered in Western Australia, have been applying AI for more than a decade now, from advanced analytics and operations, to remote decision support, to smarter maintenance, and energy efficiency across large scale assets. Today, Woodside is scaling that experience, embedding AI more deeply across its operations and the enterprise with a strong focus on governance, data quality, and human accountability.

Two words for you: technological fuel.

My guest today is Andrew Melouney, vice president for digital at Woodside Energy. Welcome, Andrew.

Andrew Melouney: Thanks, Megan. It’s great to be here.

Megan: Lovely to have you. Now, Andrew, as I said there, the energy sector has approached AI quite differently from technology or consumer businesses. Early value has emerged in operational and industrial environments, rather than consumer-facing generative AI tools. Why is that? And what differentiates the energy sector’s AI journey?

Andrew: Megan, I think it really comes down to the nature of the work we do. Energy operations and what Woodside does is very asset intensive, it’s very safety critical, and it’s highly physical. And when you think about how Woodside operates, we operate across the full value chain. We do exploration through to drilling and subsurface work, to project development, all the way through to operating assets, which are often operated in harsh and remote locations, and then global energy portfolio marketing and trading as well.

We’ve always had very large volumes of operational data coming from the equipment and the plants and the assets that we operate, and those have created really clear, quite high-value use cases for us. When you think about reliability, when you think about safety and efficiency, those are really critical things for a company like Woodside. We’ve been doing traditional AI for many years now. If you think about analytics, if you think about optimization, if you think about things like predictive models, those techniques we’ve been applying to our data sets and to our business since around 2015.

And more recently with the advent of generative AI, we’ve really found that we’ve got a pretty strong and awesome foundation to build on top of and to really solve problems in the service of improving the business. And again, whether that is keeping people safe, keeping the environments we operate in safe, or improving returns for the organization.

Megan: Fantastic. I mean you touched on it there, but how has this reality shaped your own AI strategy at Woodside? Where did you start, and where did the technology prove most impactful in those early days?

Andrew: Well, like I said, we’ve had a very long journey, in terms of understanding our operational data, recognizing the value of it, and collecting it at scale so that we can use it. And we’ve been very deliberate in that approach, Megan. We’ve really thought about where the value is and where the risks were manageable. And we’ve started looking at, in today’s world from an agentic AI perspective, we’ve started looking at the problems that were solved with traditional AI and machine learning and data science in the past. And we’ve started to think about, where can we then layer agentic AI over the top to provide an even better outcome?

For our asset intensive industry and organization, we’re looking at areas such as maintenance optimization. We’re looking at areas such as, how do we ensure our LNG plants start up reliably, consistently, and safely? And we’re considering really our frontline workforce and making sure that we’re giving people on the frontline the tools required to do their jobs. When we think about AI, we’re really thinking about, how does it support the people in the organization in terms of empowering them to make better decisions, to make faster decisions? I think over time, this has just evolved from what has been traditional analytics to now artificial intelligence and generative AI. And we’ve learned along the way that the technology is important, but it’s about aligning people, processes, and the technology together.

We’ve spent a long time not only in collecting the data and having a well-curated data set that we can build on top of, but we’ve also spent a lot of time teaching people how to work in agile ways, how to do design thinking, how to problem solve, and how to really make sure that the technology that, say, my team can bring to bear to the organization is adopted effectively and purposefully. And I think once we had that solid foundation in place from a technology perspective, from a data perspective, once we got strong trust built between our digital teams and the organization, we really saw quite a material uptick and the scaling of technology occur more broadly across the enterprise.

Megan: Fantastic. That people piece so important, isn’t it? It’s just a tool, technology, that needs to be in the right hands. And you touched on data there; industrial AI obviously depends on vast amounts of data. Can you walk us through how you’ve approached data at Woodside in a little more detail? How it’s structured and governed, and how tools like maintenance intelligence as well fit into that.

Andrew: Well, data is really foundational and fundamental to everything we do, particularly from a technology perspective. It gives us the ability to innovate at pace when we are building over the top of a strong foundation. As I said before, we’ve had the benefit of a long-term investment in our underlying operational data. I think the way we think about data is that it’s an asset for us.

And when you think about operating a facility where you’ve got sensors everywhere, you’ve got data streaming in real time, you’ve got operators needing to make decisions in real time, we have consciously made a decision over many, many years to invest in that enterprise scale data platform to make sure that it’s secure. We’ve got well-structured data assets, and we’ve got strong governance over the top of that data so that when it is used, when it’s built in a data science application or an AI agent, that we’ve got a level of trust in it that it’s going to be used responsibly. And that when it’s used, it can be trusted to give the outcome that we expect.

We have developed platforms that continuously ingest really high frequency data from the assets and from our enterprise systems. Once we’ve been able to develop solutions on top of that, parts of the business that might own the systems that collect that data, they see the value in it.

When you look at something like maintenance intelligence is a really good example of how we’ve been able to take something that we’ve been working on for a long time. Woodside does a lot of maintenance, it’s a very important part of our business, and it occurs across all of our operating assets. But we have been looking at how we do predictive analytics and predictive maintenance for a long time across that data set that we own. And something like maintenance intelligence is a solution that gives us the ability to optimize how we do that maintenance. And what it does is it analyzes historical maintenance records, alongside the performance of the equipment. And again, by having that data set well-governed and in one place, we get the ability to correlate different data sets, such as maintenance records out of SAP, alongside say equipment and performance coming from our time series data lake.

And when we build over the top of that, something like maintenance intelligence gives us the opportunity to recommend to the assets what the optimal timing for maintenance activities might be, and really give what is quite a simple aim, which is do the right work at the right time. And with something like maintenance intelligence, we have seen the opportunity, and we have the opportunity to reduce maintenance hours by up to 15% over five years on one of the assets that we’ve piloted this on. And as we’ve built out that underlying analytical model, we’re now able to put agentic AI over the top of that and provide better insights and optimize that solution more.

It really comes down to providing our asset teams and our operational teams with the right decision support capability that ensures they’re still accountable to make the decision and to ensure the right work is being done, but we are giving them the best possible opportunity to use their judgment and experience with the data that we provide to make the right decision.

Megan: Sounds like a really impactful change. Last year also marked a milestone in moving from early AI learnings to scale, using AI more deliberately as a force multiplier. What transition were you trying to make and how did you approach it?

Andrew: Well, Megan, we’ve had a philosophy for a long time in Woodside from an innovation perspective, where we really want to think big, we want to prototype small, and we want to scale fast. We want to find big opportunities that we can go after, but we want to ensure that we look at how we deploy those on a small scale first, and then provide the right learning and insight that then can scale it everywhere. Something like maintenance intelligence is a good example of that, or our Startup Advisor, where we know that we’ve got multiple plants that we need to start up. We know that we’ve got multiple assets that need to do maintenance, so we have a big, bold ambition about how we can improve and optimize that. We start with a small prototype; it might be one subsystem, it might be just a part of an asset, and then we scale it out, we learn, and we scale faster.

I think from an AI learning perspective, one of the key things we’ve learned is really the transition from moving from isolated AI solutions to a more coordinated enterprise-wide capability. If you look back maybe 18 months, two years, in our generative AI journey, we rarely started by deploying AI as broadly as we could in the organization from a personal productivity perspective. And probably being quite open in terms of the problems that we will solve, the business problems that we’ll solve with AI. That had a lot of benefits for us in terms of allowing our organization to get to know AI, get to know the capabilities, to build the trust in it.

What we’ve learned though is that we’ve needed to pivot from that to being a little bit tighter in terms of where we are going to invest our time and resources and more higher value solutions. How do we then enable and empower the rest of the organization so that they can actually effectively problem solve with technology in their domain or in their personal productivity without having to come to a central team?

When we think about that, think big, prototype small, scale fast, has been something really important for us. The transition from a more broader approach to use case development and solution development to now a narrower focus on the high value priorities. We’ve seen that paying dividends to us and allowing us to go after solutions and opportunities, things like Startup Advisor.

And so our Startup Advisor is a agentic AI solution that really aims to optimize and empower and better support our operators that sit in front of a panel and have to start up LNG plants, which are incredibly technical facilities and require really specialist skills to start up. And so our Startup Advisor is almost like a copilot that sits alongside those operators, and it gives them the ability to be able to play back previous startups. It gives them the ability to look at how the current startup is progressing, and it provides them better insights to optimize how they start up that facility. And again, starting up an LNG facility is incredibly complex.

Megan: I can imagine.

Andrew: When we think about opportunities like Startup Advisor, again, it goes back to that think big, prototype small, and scale fast. We started with a very bold vision of, how do we start up all of our LNG plants in a much more structured and optimized fashion? How do we better support our panel operators? How do we make, say, a more junior panel operator have a copilot that can help them almost like an experienced panel operator sitting next to them? And when we think about that vision and the ability then to prototype on a small scale and then scale fast, I think it’s been really successful for us.

As we scale, we’ve just naturally expanded into more agent-based solutions. Today, we’ve got around 50 AI agents in production, supporting both our operating assets and our enterprise workflows. These tools have been proven in live environments, and we have really seen the benefit of being able to shift from point solutions that maybe solve small scale problems in specific areas, to AI and agentic solutions with agency that can really work across our workflows.

We’re able to do this because we’ve standardized on the platform that we build on and we’ve got repeatable patterns. That’s been another really important learning for us, is that we don’t want to build 50 solutions in 50 different ways. We really want to be empowering our organization and our technical teams and the users of our solutions to roll them out quickly, to roll them out safely, and to do it in a patternized and platform manner.

But the last point I’ll make, Megan, from a learning perspective is that we’ve really understood that a strong governance around how AI is deployed and developed is critical for us, and it’s critical for us to go fast as well. The traditional ways of governing how we roll out different solutions or digital systems isn’t going to scale to the breadth that we need when we are thinking about AI. Being able to have a clear philosophy around how we innovate, transitioning from isolated solutions to that enterprise-wide capability, and making sure that we’ve got strong platforms with strong patterns and clear governance are the three really critical things that we’ve learned.

Megan: Such important pillars, all of them. And you’ve been working with Infosys on this journey. How has that partnership helped accelerate scaling and embedding AI across the business?

Andrew: Well, Infosys is our managed service provider, and so they play a really critical role in the operations of our core business. One of the things that I like to say is that our license to innovate is based on our license to operate. And so, for my team to be able to turn up to an operating asset or a corporate function and have the trust that’s needed to be able to innovate and reimagine and redesign how work gets done, to be able to do that, we need to make sure that our core platforms, our core systems, our applications are running really reliably, safely, and consistently every day. Having an experienced partner like Infosys looking after those core operations in partnership with our internal teams is really, really important to us.

As we move from pilots to enterprise-wide deployment, the ability to partner with someone like Infosys also gives us the ability to scale. And so being from Perth and Western Australia, while we’ve got a really strong local team in Western Australia, and we’ve also got a very strong team in some of our other operating locations, like everyone, we’re struggling to find people that can fill AI roles. Being able to partner with Infosys and have a number of different operating models at our disposal becomes really important for us. Having co-mingled teams where they are staff, they are Infosys staff, Woodside staff, and some of our other partners, really just brings diversity of thought and experience to how we solve problems.

Fundamentally, the partnership has allowed us to operate and innovate with more confidence. While Woodside always retains ownership of the strategy and where we’re going and the governance and my teams remain accountable for the outcomes, we can’t do what we do without strong partnerships like the one we have with Infosys.

Megan: Fantastic. And as AI adoption scales, you mentioned yourself, governance becomes increasingly important. How challenging has that been, and what guardrails have you put in place at Woodside?

Andrew: So, Megan, governance is really important to us, and we operate in a well-regulated environment. That means we’ve got to make really deliberate and well-reasoned decisions when we’re thinking about how we deploy technology into our organization, whether it’s artificial intelligence or anything else, for that matter. And so, governance is really central to how we approach the execution of our AI strategy at Woodside.

We’ve got maybe two or three really key things that we’ve put in place. The first one is just making sure that every AI use case goes through a structured assessment, and that’s making sure it meets our privacy controls, our cyber controls. We’re also asking the question, not just, could we do this, but should we do this? We’ve really got to bring together safety, ethics, transparency, accountability, and make sure that we make an informed decision. When an AI solution is going through that structured assessment, if there are concerns about how we might use that solution, it then goes to an AI council that’s made up of senior leaders across the organization. That council and that group really oversee some of the prioritization and risk management. That’s where we can have really strong, robust debates around, again, could we do something, should we do it, and how do we mitigate any of the risks that we might introduce here?

I think the last one, Megan, is really around lifecycle management. When you start thinking about, we’ve got 50 at the moment, but if we had 500 agents working in our organization, really amplifying the experience and the decision-making and the value creation of our staff, we really want to have an ability to manage the lifecycle of how those agents operate. We want to know, how many people are using them? What’s the efficacy and the outcome? Is there model drift? Do we need to retune or retrain? I think that’s an area where many organizations, including Woodside, are still leaning into and still figuring out the best way to do this. We can do it quite easily with 50 agents, but 500, 5,000, 50,000 becomes an opportunity for us. Again, thinking about how we partner with others, solving problems like that really present an opportunity to co-create and to co-solve with some of our partners, like with Infosys.

Megan: Fantastic. Just to close, what’s your long-term vision for AI at Woodside? How do you see this evolving over the years ahead, and what could it unlock for the sector in your view?

Andrew: So Megan, I think our ambition is really for an autonomous enterprise, where we have agents with agency that are able to really deeply interact with our core workflows. The outcome that we want to get from that is to protect our people, to protect the environments we operate in, and to be able to provide energy at a lower cost to the world. When we think about that ambition, we can really see that being applied to almost all of the areas that Woodside work in. Whether that’s from exploration through to project developments, through to operations or marketing, the scale of the opportunity in front of us and the ability for us to really change the way that work flows through the organization is really exciting.

For us, there’s three things that we have to get right in terms of being able to execute on that ambition. The first one is really thinking about how the work gets done in the organization so that we’re not just bolting AI onto an existing process, but we’re deeply thinking about how that work needs to be reimagined. We’ve also got to think about how we enable our workforce to work differently. Providing them with the skills and the tools and the ability to really harness the power of the technology that we provide.

Secondly, we’ve got to continue to move from and restrain ourselves from deploying point solutions that solve very narrow problems, to having more connected, agentic systems of systems that can interact with each other. To do that, and if we do that successfully, that’s where we really get the high value unlock from agents being able to interact with workflows and really change how the work gets done.

And lastly, Megan, it’s about how we must continue our philosophy of thinking big, prototyping small, and scaling fast.

Megan: Which is a fantastic lens to which to make all these decisions. Thank you so much, Andrew. That was Andrew Melouney, vice president for digital at Woodside Energy, whom I spoke with from Brighton in England.

That’s it for this episode of Business Lab. I’m your host, Megan Tatum. I’m a contributing editor and host for Insights, the custom publishing division of MIT Technology Review. We were founded in 1899 at the Massachusetts Institute of Technology, and you can find us in print, on the web, and at events each year around the world. For more information about us and the show, please check out our website at technologyreview.com.

This show is available wherever you get your podcasts. And if you enjoyed this episode, we hope you’ll take a moment to rate and review us. Business Lab is a production of MIT Technology Review, and this episode was produced by Giro Studios. Thanks ever so much for listening. Goodbye.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Agriculture is ready for AI, but its data isn’t

Artificial intelligence is transforming what is possible in agriculture, but industry leaders should be wary of investing in AI without first laying the groundwork. 

The use cases are promising, especially for an industry navigating volatile fertilizer costs, unpredictable weather, and margins that leave little room for error. Research shows AI-enabled predictive models can improve crop yield by 26%, reduce water use by 41%, and cut chemical usage by 33%. 

However, what AI vendors usually won’t tell you is that these solutions are only effective if you have a clean, solid data foundation. However, at Reltio, we have experience in this area, including leading technology strategy at a major agricultural distributor and building a data platform used by enterprises worldwide–we’ve seen it first hand.

What AI vendors won’t tell you 

Vendor conversations in agriculture tend to follow a familiar pattern. The pitch leads with grand promises around using AI to monitor crop health in real time, optimize irrigation, and squeeze more yield from every acre. 

The promise is compelling, but what rarely comes up is the question of whether the data foundation underneath those promises is accurate and complete. If not, there is a real and significant risk that AI will generate misleading outputs that seem authoritative but inspire action that is, at best, counterproductive. 

For instance, a yield prediction model fed inconsistent historical data will generate imprecise forecasts. Similarly, a precision irrigation system drawing on fragmented sensor data will make watering decisions that waste resources instead of saving them. 

In each case, the AI is failing because the data it was trained on was not sufficient to produce trustworthy outputs. In agriculture, every AI hallucination is a liability, and the likelihood of error is high.

Why agriculture is a uniquely challenging test case

The data landscape across a modern agricultural operation or a large distributor serving thousands of growers is extraordinarily complex.

Modern farming environments make extensive use of IoT devices and machinery. Irrigation systems are automated, tractors navigate fields autonomously, and drones capture field imagery at scale. 

However, machine data is disparate by nature. Add in external sources, including weather feeds, U.S. Department of Agriculture data, and third-party market information, and the question of how you bring all of it together into something coherent becomes a significant undertaking. 

Agricultural AI also needs to understand more than just customer attributes; it needs to understand the land: GPS coordinates, farm boundaries, field blocks, and soil variation across a single property. Where do you apply fertilizer, and at what rate, and in which specific area of the farm? Not all parts of a field are the same, and an AI system that treats them as if they are will produce recommendations that are at best imprecise and at worst damaging.

There is also a compliance dimension due to the chemicals and the responsibility involved. Operational AI in agriculture needs significantly more checks and governance than it might in a lower-stakes environment. When a flawed recommendation gets acted upon in the field, the consequences can be severe. 

What data readiness means in practice 

Data readiness is the difference between AI delivering on its promise vs. a “garbage in, garbage out” scenario. Fundamentally, being ready for AI means having a data model that accurately reflects how the business operates. 

For a company like Wilbur-Ellis, a 104-year-old, family-owned agricultural distributor, that means understanding who your customers are, which fields they farm, which inputs they need, which suppliers those inputs come from, what they paid last season, and how all of that connects to margin. That information needs to be current, consistent, and accessible across the organization, rather than locked in separate systems that were never designed to talk to each other.

Similarly, for farming operations themselves, data readiness means having a reliable, connected picture of what is happening across every field: soil health records, input application histories, yield data from previous seasons, equipment performance, and real-time sensor readings from irrigation systems.

Governance matters just as much as structure. Prices change, relationships evolve, and suppliers come and go. An AI system drawing on data that was accurate six months ago but has not been maintained will make recommendations based on a version of the business that no longer exists. 

Building the foundation that makes AI trustworthy

The good news is that the path to data readiness is feasible. It starts with a strong data model: a single, governed source of truth that connects customers, suppliers, products, pricing, orders, and margins in a way that reflects how the organization operates. 

From there, it requires data pipelines fast enough to deliver insights when decisions need to be made, governance frameworks that keep that data trustworthy over time, and security controls that ensure sensitive commercial information is accessible to the right people under the right conditions.

This is precisely the challenge that Reltio, an SAP company, was built to solve. Reltio enables companies to unify their fragmented data so AI agents and systems can operate from a complete picture of the business. Reltio builds a trusted system of context, known as the context intelligence layer, that brings all entities, relationships, rules together under one roof and makes business data easy to access and interpret.

For Wilbur-Ellis, building that trustworthy data foundation has meant being able to ask more complex questions and trust the answers, which is the precondition for any AI system to be genuinely useful.

How agriculture can drive real value from AI

The question worth asking before the next AI conversation is not whether the use case is promising. It almost certainly is. The question is whether the underlying data foundation is strong enough to make the output trustworthy. 

Agriculture has always required its leaders to make high-stakes decisions under uncertainty, and AI offers the genuine prospect of making those decisions faster and better informed. That prospect is only achievable for organizations that have done the foundational work first, and the businesses that will get the most from AI are the ones investing in that foundation now.

This content was produced by Reltio. It was not written by MIT Technology Review’s editorial staff.

Agent confidence on the technical frontier

Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mounts, executives and technology leaders are looking to agentic AI to drive the measurable financial outcomes their businesses seek.

A prime opportunity for AI agents exists in the tech function, where IT infrastructure costs are projected to grow two to three times by 2030, even as budgets remain unchanged, according to McKinsey. And in the last 18 months, tech teams—the engineers, developers, architects, and other practitioners who are building, deploying, and continually improving their organizations’ infrastructure and applications—are clearly putting agents to work.

The ultimate promise of agents is not only to automate tasks but to manage and coordinate entire workflows, pursuing business goals in a way that allows humans and agents to work together. Given the risks involved in automated decision-making, teams cannot delegate the work that agents do without confidence that they are fully capable of performing the task and that it will do so in a safe, reliable, and secure manner.

Among technology experts, our research shows that teams are exceedingly confident about using agentic AI across a significant amount of AI, data, and cloud tasks.

Where agent readiness drops is largely due to a lack of business context being supplied to agentic systems. The more complex the task, the more reasoning capability an agent requires and the greater its need for business context. Such context-generation capabilities for agents are still at an early stage of development, especially in situations where enterprise data is difficult to wrangle and connect into the agent lifecycle at the speed and quality in which developers and executives need it. Human oversight is a key factor of success in deploying agentic AI.

Knowing that tech teams are in a pivotal position to lead this transformation, the experts we interviewed expect agent confidence to accelerate as experience with agents deepens and business environments mature. “As we design agents to operate within the same operational boundaries, identity systems, and governance models that teams already use, they start to behave more like the systems organizations already trust,” says Jeremy Winter, corporate vice president and chief product officer at Microsoft Azure Platform.

This report, based on a survey of 300 global technology experts, ranks 101 tasks across AI, data, and cloud workflows based on respondents’ confidence in agents acting on their behalf. It also examines how technology teams view the opportunities and challenges related to agentic AI, along with the potential for the technology to enhance their careers.

Key findings from the report include:

Confidence in agents is surging for measurable tasks and growing in areas of complex judgment. Technology experts overwhelmingly believe agents help with everyday work including streamlining processes, improving performance, and reducing repetitive tasks. Confidence is highest for processes like generating reports and boilerplate code, and there is clear opportunity where tasks involve multistep workflows and advanced reasoning to make decisions.

Data workflows are the breakthrough domain. Tech teams trust agents most where structure can provide a reliable foundation for decisions. This includes areas such as data quality monitoring, visualization anomaly detection, real-time data stream monitoring, and data profiling. This is where domain experts closest to the point of data generation can provide context to allow agents to act and deliver trusted outcomes.

Download the full report.

Read the Microsoft Cloud blog by Amanda Silver, corporate vice president of Microsoft 365 Core and Work IQ, which underscores the importance of keeping humans in the loop and how systems thinking advances careers. And for a deeper dive into data workflows as a breakthrough use case for agents, check out the Fabric blog to hear from Kim Manis, corporate vice president of Product for Microsoft Fabric.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.




Repositioning retail for the AI era

Artificial intelligence is rapidly reshaping retail, but not in the ways consumers might immediately notice. The biggest transformation may not be flashy virtual try-ons or chatbot shopping assistants, but in how decisions are made behind the scenes: how products surface in search results, how inventory moves through supply chains, how engineers ship code faster, and how retailers respond to customer behavior in real time. As legacy retailers navigate a fragmented and hyper-competitive landscape, AI is becoming an operating philosophy.


At Macy’s, that philosophy is more often defined by what senior director of engineering Murali Murugan describes as an “AI-first” approach. “AI first isn’t about adding intelligence on top,” Murugan says. “It’s about redesigning how decisions happen so the business moves faster and every experience feels more relevant by default.” Rather than layering AI onto existing workflows, Macy’s is embedding intelligence directly into systems that include personalization, search, operational planning, and software development itself.

The company’s strategy is reflective of a larger shift taking place across retail: moving from isolated AI pilots toward integrated systems designed to compress, as Murugan puts it, “the gap between the signal and the action.” Early efforts focused on narrow, high-impact use cases like search recommendations and customer engagement, where measurable gains in conversion and reduced friction quickly built internal momentum. “Once we established the quick wins, scaling was a business decision, not a technology debate anymore,” he says.

That momentum is now extending into conversational commerce through tools like Ask Macy’s, an AI-powered shopping assistant designed to act more like a personal stylist than a traditional search bar. Whether for a prom, a vacation, or a last-minute event, customers can describe what they need conversationally and receive curated recommendations informed by past purchases, preferences, and context.

Still, the company sees AI as more of an invisible layer augmenting human judgment than a replacement for it. The long-term vision is retail that feels increasingly seamless, adaptive, and personalized, powered by systems customers may never even notice are there.

“The real transformation in this all comes from continuous improvement,” Murugan says. “It’s about learning from the mistakes, quickly adapting to the newer technology standards that are coming into play, timing, and execution which compound into a meaningfully better customer experience.” 

This webcast is produced in partnership with Infosys.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

The emergence of the web data infrastructure layer for AI

AI is booming. New use cases are emerging each day. To capitalize on the technology’s potential, enterprises require data at scale. In many cases, though, the relevant information is blocked or unstructured, which limits its use by AI models. 

To understand this challenge, consider the foundation of the web itself. The web was not designed for the automated discovery and retrieval that new AI applications demand. Overcoming this inherent design constraint requires infrastructure.

The next frontier in AI may depend on a new web data infrastructure layer that can enable models to discover and map this ever-expanding digital realm. This layer must be able to navigate hundreds of millions of existing web domains and billions of new URLs created each week, delivering real-time information and overcoming technical barriers.

“The data suggests there’s far more data out there,” says Or Lenchner, CEO of Bright Data, a web data collection platform. “Think of the universe: It’s out there, but you don’t know what you don’t know.”

Enabling access to fresh, relevant, and trustworthy data

While early AI breakthroughs were driven by scaling training data and model size, organizations are now encountering a fundamental bottleneck: They need to keep pace with the dynamic, unstructured, and constantly evolving nature of web data in order to ground outputs in current and verifiable information. AI performance increasingly depends not just on model architecture but on a system’s compute, networking, retrieval, and data engineering capabilities—that is, the system’s ability to quickly and reliably retrieve data that is fresh, relevant, and trustworthy.

Traditional model training relies on snapshots of information collected at a particular point in time. Training AI on such static data is no longer sufficient. To track fluctuations such as competitor pricing, consumer sentiment, and market trends, companies need a constant feed of new information, pulling data in real time along with relevant context. Their infrastructure must therefore be able to handle millions of simultaneous interactions across websites that vary by geography, language, format, and access rules.

“If it can’t retrieve real-time information, it lacks context,” Lenchner says. “In a business setting, that’s not acceptable anymore. Stale answers lead to bad decisions and disappointed consumers.”

Speed is not merely a matter of convenience; it’s a matter of necessity. Today’s organizations operate in environments where prices, inventory, markets, security threats, and customer behavior change continuously. Delayed data retrieval can reduce the usefulness of an otherwise sophisticated model.

Using live, high-quality web data can also reduce AI hallucinations because the model has a more relevant knowledge base. This builds user trust. In fact, one survey found that 56% of AI practitioners said businesses need access to real-time web data to improve trust in AI outputs. To ensure the model runs efficiently and effectively, the information must also be pared down to the appropriate essentials. 

Despite the introduction of retrieval-augmented generation (RAG), where models pull in external data at the moment of a query, many AI systems still struggle to deliver outputs that are current, contextually relevant, and trustworthy in operational settings. According to Gartner, 60% of AI projects that are not supported by AI-ready data—accurate, structured, organized, and contextualized—will be abandoned by the end of the year. 

This is because large-scale retrieval alone does not solve the problem. As Lenchner puts it, “You need to retrieve data at scale, but also in real time. Latency becomes an issue because of the end user who is waiting for the output.” 

Accessing fresh, AI-ready data at scale introduces technical and structural challenges. In practice, many enterprise systems combine public web retrieval with APIs, licensed datasets, and proprietary internal data in their AI applications. Integrating these fragmented sources into a timely and usable knowledge layer requires specialized capabilities. Some research has found that 97% of AI organizations depend on real-time web data infrastructure, but 90% feel boxed in by various restrictions. Companies are increasingly developing technical approaches to navigate these constraints.

Lenchner draws this metaphor: “Think of the trained model as intelligence and relevant data as knowledge. A powerful intelligence layer sitting on top of a hollow knowledge layer is like a genius who knows nothing—useless in practice. Intelligence and knowledge have to come together.”

The promise of new infrastructure

A new layer of web data infrastructure can address this developing need for stronger AI inputs by enabling discovery of data, real-time access, and tailoring to a specific context. As Lechner describes it, “It’s all about collecting data at scale, super-low latency, without being blocked.”

Rather than relying on increased computing power, this type of platform emulates human browsing behavior to access available content and transform raw code into structured data feeds. It can work with websites that might not interact with traditional scraping tools, such as those heavy in JavaScript, or with aggressive antibot software. 

As Lenchner explains, “It’s basically having infrastructure that can mimic a web user with identifying information—IP address, location, and 1,000 more parameters. And at scale. Think of doing that 80 billion times a day for millions of websites. And every single time, you are looking exactly as the website expects you to look.”

Of course, continuous retrieval introduces new data governance challenges. To address them, platforms can enforce strict compliance protocols aligned with global privacy frameworks, such as the EU’s General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). They can also be limited to openly accessible, public information, avoiding paywalls or private logins. Any networks used can be vetted and consent-based, and incentives can be provided to owners of IP addresses. In this way, systems can be designed to comply with tightening regulation.

Such complex capabilities do not come easy. “When this is critical infrastructure for a company,” Lenchner says, “doing it in-house becomes a full-time engineering problem that competes with the actual AI work.” Addressing this complexity requires organizations to commit significant resources, leading many to seek specialized platforms designed specifically for data retrieval, orchestration, and observability.

Infrastructure for the real world

Real-time data retrieval is changing what AI systems can do inside organizations. For example, a retail company can use public information to enable a dynamic pricing engine, and global brands can track trademark infringements. 

As the ecosystem matures, organizations that invest in this emerging data infrastructure layer will be better positioned to build AI systems that are more responsive, reliable, and aligned with real-world conditions—AI systems that can continuously adapt using current web data. Over time, the distinction between AI models and the infrastructure that feeds them may even begin to disappear.

As Lenchner says, “The world is changing. And everything that is happening in the world is being uploaded to the public web. The amount of new data that is being generated is growing and accelerating.”

To learn more from Bright Data, read the Data for AI 2026 report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Learning to lead in a hybrid human-AI enterprise

As adoption of AI agents looks set to surge by as much as 300% in the next two years, leadership teams are carefully considering the implications of a hybrid human-AI workforce. 

Unlike existing enterprise-level automation that relies on manual input, AI agents are capable of autonomously coordinating complex tasks, interacting with multiple tools and environments across an organization. In early applications that center on customer service, HR, and sales, adoption of agentic AI has led to productivity gains of 30-50%

Their autonomy positions agents more as collaborators than tools, working side-by-side with human employees in blended teams that look poised to upend traditional workplace dynamics. 

More than three-quarters of HR leaders believe that the deployment of AI agents will transform existing workplace norms, driving a complete reappraisal of how roles and responsibilities are distributed, how skills are prioritized, and how workplace culture is shaped.

Though many admit they’re in the early or preparatory phase of this shift, 86% of chief HR officers predict that navigating digital labor shaped by agentic AI will be a central component of their role in the years ahead.

Fluency in the change management aspect of agentic AI adoption will be a crucial differentiator when it comes to unlocking the full potential of the technology going forward, believes Ateet Jayaswal, chief culture and employee experience officer at Wipro, a leading technology services and consulting company. This moment is one that he says, “calls for a mindset shift in how HR leaders would enable their organizations.”

Redeploying roles to enable higher-value work

As AI agents assume ownership of more complex and integral tasks, the distribution of roles and responsibilities within an organization will undergo significant change. It’s estimated that three-quarters of current roles will require redesign, reskilling, or redeployment by 2030 as a result of agentic AI. 

For leadership, this shift should be about reskilling employees toward higher-value work in order to optimize the potential of an agent-human hybrid workforce, says Jayaswal. 

For example, Wipro is a complex organization of 240,000 employees across 65 countries. It previously had multiple policies, documents, and knowledge fragmented across different systems, which delayed response to employee queries. 

But the company has recently integrated a custom agentic AI assistant—an agent co-created in partnership with enterprise agentic AI platform Ema Unlimited—that can swiftly navigate this complex system, assuming responsibility for 50 HR tasks that had previously fallen to human employees. With the help of an AI agent, average response time to queries has lowered from 48 hours to five seconds. 

Human employees have more time to focus on work “that requires a creative and imaginative mind and cross-functional collaboration, leveraging diverse ideas and thoughts to problem-solve,” says Jayaswal. The AI agent, meanwhile, handles rote administrative tasks like sorting timesheets or helping employees navigate policies and take actions in the flow of work. 

When reallocating employee responsibilities, though, it is imperative that humans remain in the loop, Jayaswal caveats. When agentic AI is incorporated into enterprise technology, it must work with sensitive and personal data and therefore needs even more stringent guardrails and constraints than consumer applications. “When you expose an AI agent to organizational data, when you integrate it into multiple enterprise systems, then pathways around the AI agent become extremely important,” he says. “It’s an evolving space that leadership needs to have front-of-mind.” Governance should include robust data privacy rules and the establishment of governance layers, such as an AI council, he suggests.  

At a fundamental level, the adoption of AI agents will force a re-evaluation of human roles, believes Jayaswal. Rather than employees primarily performing repetitive tasks or troubleshooting, a significant proportion of their time will shift to designing, teaching, and optimizing an AI agent that can do this work for them with far greater speed and predictability and without the agent getting bored. 

“The nature of your job changes from being the hero who comes in to solve the problem to designing the hero who can solve the problem,” he summarizes. “The individuals who I have seen thrive in this environment are the ones who make this shift.”

An evolving employee skillset

Just as roles and responsibilities will be reconfigured to reflect the input of AI agents, the core skills of human employees will be reprioritized. More than four in five HR leaders say they’re planning to reskill workers to become more competitive in a market shaped by AI agents. 

Technical skills will be increasingly important. Leading employers such as Salesforce, Danone, and Walmart are already rolling out dedicated AI and digital skills programs that aim to equip everyone from frontline workers to C-suite executives with a baseline level of AI literacy in response to the pervasiveness of the technology. 

But desirable soft skills will also evolve, Jayaswal points out. Employees who assign tasks to an AI agent need to plainly articulate what modular steps may be needed to accomplish a task, what the desired outcome should be, and what parameters or guardrails need to be in place to ensure the agent doesn’t access or share confidential data. 

As HR executives adapt to a blended workforce, three skills are emerging as top priorities during recruitment, according to a recent survey: relationship building, like forging constructive partnerships and account management; collaboration; and adaptability. 

Maintaining a healthy workplace culture

In freeing up human employees to focus on higher-value tasks, the hope is that AI agents can elevate the employee experience, deepening fulfilment and satisfaction in the workplace. 

“At Wipro, our vision is to improve the life of Wiproites,” says Jayaswal. “We are taking away non-value added work by embracing modern ways of collaborating, engaging, and transacting, leaving associates with higher order work content.” 

But leadership teams embracing agentic AI will also need to plan for the new pressures and stressors that the technology can place on a workforce. 

There is already confusion and knowledge gaps, with 73% of HR leaders reporting their employees don’t yet understand how digital labor will impact their work. Many organizations have opted to define AI agents as teammates or colleagues on org charts, but new research says this could erode trust and a sense of professional identity. It also raises new questions around accountability and ownership. 

The role of management in addressing these concerns is critical, says Jayaswal. To maintain healthy dynamics, managers need to become skilled at orchestrating blended systems, splitting their focus between supervising AI agents and motivating human employees as they also build and supervise AI agents.

Upgrading employee well-being programs will be a core part of maintaining a robust workplace culture. “As there are more interactions with AI agents, you are losing some of the human touch that was provided by service delivery partners or leaders, or often even by colleagues and peers,” Jayaswal says. Employee services that encourage social connection and empathetic communication may help teams navigate this. 

A breakneck transformation

Agentic AI looks set to scale at breakneck speed across many enterprises, and it will significantly transform how these organizations operate. 

Carefully considering and deciding how to adapt to this newly blended workforce is now a top priority for leadership teams. Reviewing and refining organizational strategies is essential for optimizing both technological gains and the employee experience.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Rehumanizing global health care with agentic AI

The global health care sector is under increasing strain. 

Decades of chronic underinvestment and constraints in recruitment have coincided with a surge in demand for services for aging populations. Gaps in provision are already taking a toll, with fragmented access to care and high rates of stress and burnout among staff. And it’s getting worse. The World Health Organization has warned that current shortfalls will increase to 11 million workers by 2030. 

In their urgent hunt for a solution, many health-care providers are now pinning their hopes on agentic AI, with more than two-thirds (68%) having already adopted AI agents into their workforce, according to KPMG. 

The technology is being deployed to automate complex back-office processes, collaborate with medical teams, and even triage patients, all in a bid to reduce the cognitive load on clinicians and improve quality of care for patients as the supply of human health-care workers dwindles.

A different type of digitalization 

Until now, the benefits of digitalization within health care have been limited. 

Many staff have blamed slow or outdated technology for adding to the administrative burden rather than alleviating it. For example, U.S. patient data was migrated to electronic health records (EHRs) in the early 2000s, but this data remains fragmented and reliant on manual inputs. 

New telehealth services and digital care tools, like remote monitors, have had similar shortcomings, says Ashis Barad, MD, chief digital and technology officer at Hospital for Special Surgery (HSS), an academic medical center in New York that focuses on musculoskeletal health. Both technologies have helped improve access to health care by removing geographical barriers, he says, but they’ve failed to replicate the quality of in-person care or win trust from patients. 

Agentic AI is different from these existing technologies, he insists. 

Rather than relying on manual inputs or defaulting to human workers for any case that sits slightly outside a rigid framework, AI agents can handle nuanced, complex scenarios. They can make autonomous decisions, retrieve information from expert clinical sources, and iterate over time, freeing clinicians to focus on higher-level patient care. As Dr. Barad puts it: “Agentic AI takes your workflow and collapses it, augments it, supercharges it, and makes it more performant.” 

At HSS, AI agents have already been deployed in multiple areas. They handle complex backend processes, such as insurance claims that previously took several weeks to complete and involved both HSS staff and a third-party contractor to handle the volume. Now, says Dr. Barad, AI agents complete 1,100 claims per month. They’ve reduced the appeals stage from 45 minutes to five and improved the success rate of those appeals from 65% to 100% in the nine months since implementation. HSS now handles all claims in-house. 

Building on that success, HSS is now deploying AI agents in non-clinical patient-facing settings with an AI scheduling and triage service, as part of a collaboration with enterprise agentic AI developer Ema Unlimited. The service is accessible 24/7 via web, text, or phone. It uses conversational AI to ask patients clarifying questions about their condition and then books appointments with the most appropriate clinician, factoring in location, insurance coverage, and physician availability. “It completes the whole loop,” says Dr. Barad. The AI agent is trained on “all of our context, all of our rules, and all of our knowledge base,” he adds, providing patients with streamlined access to highly specialist knowledge from world-leading surgeons.

Given the high-stakes decisions delegated to AI agents, the triage service has built-in safeguards—sensitive, complex, or uncertain scenarios are escalated to human specialists. Every decision made by the AI agent is auditable and human staff can step in at any point. Patient data is kept secure and the system is trained on all HSS protocols, policies, and care pathways. By keeping humans in the loop, Ema says its technology strikes the balance between efficient automation, patient-first safety, and human-informed decision making. 

As the technology becomes more prolific, it will be incumbent on providers to ensure they have these sorts of guardrails embedded into systems, says Dr. Barad. At HSS all decisions around the technology are filtered through an AI subcommittee that Dr. Barad co-chairs alongside a senior nursing executive. AI agents that may touch on patient care will be scrutinized with far more rigor than, say, backend processes, he explains.

AI agents prompt systems-level change

For example, Dr. Barad has plans to create a dedicated AI lab at the HSS main campus in New York City—a move that aims to democratize access to the technology across the organization. It will be open to all staff looking to understand or build AI agents, he explains, with informative classes and one-on-one training. “We’re getting agentic AI into everybody’s hands,” he says. This echoes research by Deloitte, which found that leading agentic AI adopters in health care were far more likely to have opted for multiagent solutions, redesigning end-to-end workflows rather than sticking to narrow solutions or individual use cases.

The key, it appears, is to integrate AI agents across the entire enterprise, treating them as a general-purpose technology. As Dr. Barad puts it: “It’s wrong to think of agentic AI in use cases… It’s a general-purpose technology, analogous to electricity.”

In practice, this means health-care providers need to set the right foundation to achieve value with agentic AI. This includes creating a unified data strategy, one that integrates fragmented data sources across an organization to create a single, comprehensive source of truth. In health care, data is often split across multiple departments and providers, each with their own legacy IT system.

In systems that rely on fragmented data sources, metrics often lack standardized definitions too. For example, Dr. Barad says that each hospital he’s worked in has had a slightly different definition for “time to start surgery,” a metric commonly used to gauge operating room efficiency. This level of fragmentation impedes AI agents from retrieving information from different sources or applications and assimilating the tacit knowledge that differentiates them from other technologies.

By creating greater interoperability of data at HSS, patient-facing AI agents can draw from a patient’s clinical care history and existing recommendations from their clinician, combine this information with current symptoms, and decide whether a situation requires escalation before notifying the correct specialist and informing the patient. 

Building better outcomes

For Dr. Barad, the potential for AI agents to overhaul health care and alleviate the current pressures on resources, access, and patient care is huge. 

He envisions a future in which 90% of non-clinical health-care tasks could be administered by AI agents, freeing clinicians up for what he calls white-glove work, meaning the most complex, specialized, and sensitive cases.

Most health-care providers seem equally optimistic. According to research by KPMG, 84% of providers are already comfortable handing decision making about specific processes over to AI agents.

“We’re spending so much time on keyboards and computers right now that we’re actually not doing what we should be doing,” says Dr. Barad. “This is going to rehumanize health care.”

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Rethinking organizational design in the age of agentic AI

Amid rapidly growing adoption of enterprise-level AI agents, there’s a disconnect emerging between ambition and execution. 

Although 85% of organizations say they want to be agentic within the next three years, 76% say their current operations and infrastructure can’t support that change. They cite a lack of readiness across people, processes, and workflows. 

The sticky tape problem

The challenge is that many organisations are often layering AI agents onto existing operations, rather than reimagine the operating model and how work will need to be rewired, explains Prasun Shah, global CTO for workforce consulting and chief AI officer at PwC UK Consulting. “They’re embedding AI employees into what is a human operating model,” layering on AI agents to existing workplace structures when “this is like adding sticky tapes to parts of an operating model that is breaking.”

Doing so may be preventing organizations from unlocking the full value agentic AI offers, creating circumstances where disillusionment can quickly creep in. That full value lies in agents’ capacity to execute entire workflows with limited human input. They can coordinate complex tasks, make independent decisions, adjust to changing conditions, and iterate performance. 

In early proving grounds that span customer service, HR, and sales, it’s already estimated that AI agents could accelerate business processes by as much as 30% to 50% and low-value work time by 25% to 40% when deployed at scale. But with this capability comes greater complexity and the need for an enterprise-wide change.

Growing the AI vocabulary 

Enterprise agentic AI platform Ema describes this change as agentic business transformation (ABT), a term it coined last year in partnership with HFS Research, in an attempt to plug what it sees as a gap in the existing lexicon about AI agents, and to provide enterprises with a new framework with which to think about their own adoption of the technology. 

“None of the existing vocabulary captures the full scope of the change,” explains Ema CEO and founder Surojit Chatterjee. “Digital transformation was about moving from paper to software. AI transformation was about adding artificial intelligence to existing processes. Co-pilot is about AI assisting in various human tasks. But ABT is something categorically different: It’s the integration of AI agents into the fabric of the organization.” 

For Shah, the dedicated term (ABT) “helps drive the need to redesign an organization in its entirety: its operating model, its workflows, decision rights, and performance management systems.” He emphasizes that “everything that’s needed to ensure those agents are actually active participants in value creation, rather than just point tools or productivity aids.”

According to Ema, ABT encompasses three core pillars: an organization’s technology stack, its workforce, and the metrics used for success. 

AI agents as connective tissue

The first pillar of ABT is the technology stack. “Your existing tech stack was designed for human-operated, application-centric workflows,” says Chatterjee. “It needs to be reconsidered when the actor is an AI agent operating at machine speed across multiple systems simultaneously.”

 As AI agents are integrated into an organization, enterprises will need to pivot from a set of linear processes and steps, to rewiring work in a very different way, explains Shah. That’s because the value in AI agents isn’t as another layer in an existing technology stack but as a connective tissue, he explains, moving between or across layers to coordinate a high-level task or retrieve and interpret data from multiple discrete applications. AI agents can create “a true competitive differentiation for an enterprise” by making decisions based on this capacity to contextualize, he says. “That is where the next battleground will be.”

To build this connective tissue, leaders need to adapt their technology stack to surface higher quality decisions from AI agents, prioritizing access to multiple datasets and applications simultaneously to develop tacit knowledge. “Organizations that make this architectural shift become genuinely more adaptive,” says Chatterjee. “When a new business requirement emerges, you don’t wait six months for a software vendor to build a feature. You configure an AI employee using natural language and connect it to the systems it needs. The time from business to production workflow drops from months to days.”

The workforce, redesigned

As AI agents are deployed for more use cases, enterprise leaders must consider what this means for dynamics across their workforce, the second pillar of ABT.

Workforce structures today deviate little from the hierarchical model of the early days of industrialization. To maximize efficiency and scale, processes are standardized, tasks are clearly delineated between strategic business units (SBUs), and employees progress up through an organization based on their capacity to optimize output from teams below them. But with AI agents that can execute, coordinate, and optimize tasks—often without managerial coordination—the lines of that established hierarchy become blurred.

In a workforce that blends AI agents and human employees, managers will be freed up from many execution-based tasks but take on new responsibilities associated with managing hybrid teams. Managers “will need to be able to manage issues around trust, explainability, psychological safety, and even status dynamics” to navigate new tensions that could arise in a hybrid workforce, says Shah.

The impact of agentic AI on existing workforce structures goes far beyond the management layer, too. McKinsey predicts that by 2030, three-quarters of current jobs will require redesign, upskilling, or redeployment, and organizations will need to act swiftly to amend recruitment, retention, and remuneration. 

From output to outcome

Success metrics are the third and final pillar of ABT. 

As AI agents assume greater ownership of core enterprise processes, taking on collaborative roles alongside human employees, traditional workforce metrics that focus on activity or output—such as calls handled or reports filed—no longer make sense. 

“When you add AI employees into the workforce, activity metrics become meaningless or actively misleading,” says Chatterjee. “An AI employee can handle a thousand customer interactions in the time it takes a human to handle ten. If you measure success by interactions handled, you’ll conclude the AI is working brilliantly while missing whether any of those interactions actually drove customer satisfaction, retention, or revenue.” To correct this, enterprises must develop a new set of metrics that focus on outcome rather than output. That is, metrics on the broader benefits or changes achieved, rather than individual deliverables. 

For example, when one of Ema’s large enterprise customers overhauled its own metrics, switching from tool metrics like cost per query and AI accuracy, to outcomes like the percentage of contracts reviewed without human escalation, the measured ROI from agentic AI tripled within two quarters. The changes meant “this customer stopped building point solutions in high-volume, low-complexity workflows and started deploying AI employees where the outcome value was highest,” says Chatterjee.

Integrating new metrics may also require a complete reconfiguration of reward and talent management processes, as well as accountability and ownership within organizations, points out Shah. In human-AI teams, for example, although ethical and fiduciary responsibilities will likely remain with human employees, operational accountability will become significantly more diffused to reflect the systemic role of AI agents.

This change will raise new questions that senior leadership teams will need to wrestle with, Shah adds. They’ll need to consider: Who is accountable when an AI employee makes a mistake? What happens when AI and humans disagree? What guardrails should be erected to safeguard customers? 

Laying the groundwork for systems-level change

Systems-level change is gradual. These are complex lines of inquiry that experts continue to grapple with. But in kickstarting internal dialogue about the core pillars of ABT—the workforce, the technology stack, and the metrics by which success can be gauged—leaders can lay the groundwork for an enterprise better poised to embrace AI agents at a systems level and start to close the gap between their ambition and execution. 

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

DAIMON Robotics Wants to Give Robot Hands a Sense of Touch

4 May 2026 at 11:08


This article is brought to you by DAIMON Robotics.

This April, Hong Kong-based DAIMON Robotics has released Daimon-Infinity, which it describes as the largest omni-modal robotic dataset for physical AI, featuring high resolution tactile sensing and spanning a wide range of tasks from folding laundry at home to manufacturing on factory assembly lines. The project is supported by collaborative efforts of partners across China and the globe, including Google DeepMind, Northwestern University, and the National University of Singapore.

The move signals a key strategic initiative for DAIMON, a two-and-a-half-year-old company known for its advanced tactile sensor hardware, most notably a monochromatic, vision-based tactile sensor that packs over 110,000 effective sensing units into a fingertip-sized module. Drawing on its high-resolution tactile sensing technology and a distributed out-of-lab collection network capable of generating millions of hours of data annually, DAIMON is building large-scale robot manipulation datasets that include vast amounts of tactile sensing data. To accelerate the real-world deployment of embodied AI, the company has also open-sourced 10,000 hours of its data.

Person in navy suit and blue striped tie against a blue studio backdrop Prof. Michael Yu Wang, co-founder and chief scientist at DAIMON Robotics, has pioneered Vision-Tactile-Language-Action (VTLA) architecture, elevating the tactile to a modality on par with vision.DAIMON Robotics

Behind the strategy is Prof. Michael Yu Wang, DAIMON’s co-founder and chief scientist. Prof. Wang earned his PhD at Carnegie Mellon — studying manipulation under Matt Mason — and went on to found the Robotics Institute at the Hong Kong University of Science and Technology. An IEEE Fellow and former Editor-in-Chief of IEEE Transactions on Automation Science and Engineering, he has spent roughly four decades in the field. His objective is to address the missing “insensitivity” of robot manipulation, which practically relies on the dominant Vision-Language-Action (VLA) model. He and his team have pioneered Vision-Tactile-Language-Action (VTLA) architecture, elevating the tactile to a modality on par with vision.

We spoke with Prof. Wang about how tactile feedback aims to change dexterous manipulation, how the dataset initiative is foreseen to improve our understanding of robotic hands in natural environments, and where — from hotels to convenience stores in China — he sees touch-enabled robots making their first real-world inroads.

Daimon-Infinity is the world’s largest omni-modal dataset for Physical AI, featuring million-hour scale multimodal data, ultra-high-res tactile feedback, data from 80+ real scenarios and 2,000+ human skills, and more.DAIMON Robotics

The Dataset Initiative

This month, DAIMON Robotics released the largest and most comprehensive robotic manipulation dataset with multiple leading academic institutions and enterprises. Why releasing the dataset now, rather than continuing to focus on product development? What impact will this have on the embodied intelligence industry?

DAIMON Robotics has been around for almost two and a half years. We have been committed to developing high-resolution, multimodal tactile sensing devices to perceive the interaction between a robot’s hand (particularly its fingertips) and objects. Our devices have become quite robust. They are now accepted and used by a large segment of users, including academic and research institutes as well as leading humanoid robotics companies.

As embodied AI continues to advance, the critical role of data has been clearer. Data scarcity remains a primary bottleneck in robot learning, particularly the lack of physical interaction data, which is essential for robots to operate effectively in the real world. Consequently, data quality, reliability, and cost have become major concerns in both research and commercial development.

This is exactly where DAIMON excels. Our vision-based tactile technology captures high-quality, multimodal tactile data. Beyond basic contact forces, it records deformation, slip and friction, material properties and surface textures — enabling a comprehensive reconstruction of physical interactions. Building on our expertise in multimodal fusion, we have developed a robust data processing pipeline that seamlessly integrates tactile feedback with vision, motion trajectories, and natural language, transforming raw inputs into training-ready dataset for machine learning models.

Recognizing the industry-wide data gap, we view large-scale data collection not only as our unique competitive advantage, but as a responsibility to the broader community.

By building and open-sourcing the dataset, we aim to provide the high-quality “fuel” needed to power embodied AI, ultimately accelerating the real-world deployment of general-purpose robotic foundation models.

The robotics industry is highly competitive, and many teams have chosen to focus on data. DAIMON is releasing a large and highly comprehensive cross-embodiment, vision-based tactile multimodal robotic manipulation dataset. How were you able to achieve this?

We have a dedicated in-house team focused on expanding our capabilities, including building hardware devices and developing our own large-scale model. Although we are a relatively small company, our core tactile sensing technology and innovative data collection paradigm enable us to build large-scale dataset.

Our approach is to broaden our offering. We have built the world’s largest distributed out-of-lab data collection network. Rather than relying on centralized data factories, this lightweight and scalable system allows data to be gathered across diverse real-world environments, enabling us to generate millions of hours of data per year.

“To drive the advancement of the entire embodied AI field, we have open-sourced 10,000 hours of the dataset for the broader community.” —Prof. Michael Yu Wang, DAIMON Robotics

This dataset is being jointly developed with several institutions worldwide. What roles did they play in its development, and how will the dataset benefit their research and products?

Besides China based teams, our partners include leading research groups from universities, such as Northwestern University and the National University of Singapore, as well as top global enterprises like Google DeepMind and China Mobile. Their decision to partner with DAIMON is a strong testament to the value of our tactile-rich dataset.

Among the companies involved there are some that have already built their own models but are now incorporating tactile information. By deploying our data collection devices across research, manufacturing and other real-world scenarios, they help us to gather highly practical, application-driven data. In turn, our partners leverage the data to train models tailored to their specific use cases. Furthermore, to drive the advancement of the entire embodied AI field, we have open-sourced 10,000 hours of the dataset for the broader community.

Robotic gripper delicately holding a cracked eggshell in a dimly lit roomEquipped with Daimon’s visuotactile sensor, the gripper delicately senses contact and precisely controls force to pick up a fragile eggshell.Daimon Robotics

From VLA to VTLA: Why Tactile Sensing Changes the Equation

The mainstream paradigm in robotics is currently the Vision-Language-Action (VLA) model, but your team has proposed a Vision-Tactile-Language-Action (VTLA) model. Why is it necessary to incorporate tactile sensing? What does it enable robots to achieve, and which tasks are likely to fail without tactile feedback?

Over these years of working to make generalist robots capable of performing manipulation tasks, especially dexterous manipulation — not just power grasping or holding an object, but manipulating objects and using tools to impart forces and motion onto parts — we see these robots being used in household as well as industrial assembly settings.

It is well established that tactile information is essential for providing feedback about contact states so that robots can guide their hands and fingers to perform reliable manipulation. Without tactile sensing, robots are severely limited. They struggle to locate objects in dark environments, and without slip detection, they can easily drop fragile items like glass. Furthermore, the inability to precisely control force often leads to failed manipulation tasks or, in severe cases, physical damage. Naturally, the VLA approach needs to be enhanced to incorporate tactile information. We expanded the VLA framework to incorporate tactile data, creating the VTLA model.

An additional benefit of our tactile sensor is that it is vision-based: We capture visual images of the deformation on the fingertip surface. We capture multiple images in a time sequence that encodes contact information, from which we can infer forces and other contact states. This aligns well with the visual framework that VLA is based upon. Having tactile information in a visual image format makes it naturally suitable for integration into the VLA framework, transforming it into a VTLA system. That is the key advantage: Vision-based tactile sensors provide very high resolution at the pixel level, and this data can be incorporated into the framework, whether it is an end-to-end model or another type of architecture.

Close-up of a vision-based tactile sensor with 110,000 sensing units, resembling a smartwatch screen glowing with colorful digital static in the darkDAIMON has been known for its vision-based tactile sensors that can pack over 110,000 effective sensing units.DAIMON Robotics

The Technology: Monochromatic Vision-based Tactile Sensing

You and your team have spent many years deeply engaged in vision-based tactile sensing and have developed the world’s first monochromatic vision-based tactile sensing technology. Why did you choose this technical path?

Once we started investigating tactile sensors, we understood our needs. We wanted sensors that closely mimic what we have under our fingertip skin. Physiological studies have well documented the capabilities humans have at their fingertips — knowing what we touch, what kind of material it is, how forces are distributed, and whether it is moving into the right position as our brain controls our hands. We knew that replicating these capabilities on a robot hand’s fingertips would help considerably.

When we surveyed existing technologies, we found many types, including vision-based tactile sensors with tri-color optics and other simpler designs. We decided to integrate the best of these into an engineering-robust solution that works well without being overly complicated, keeping cost, reliability, and sensitivity within a satisfactory range, thus ultimately developing a monochromatic vision-based tactile sensing technique. This is fundamentally an engineering approach rather than a purely scientific one, since a great deal of foundational research already existed. With the growing realization of the necessity of tactile data, all of this will advance hand in hand.

Daimon tactile sensor showing force, geometry, material, and contact data visualizations.DAIMON vision-based tactile sensor captures high-quality, multimodal tactile data.DAIMON Robotics

Last year, DAIMON launched a multi-dimensional, high-resolution, high-frequency vision-based tactile sensor. Compared with traditional tactile sensors, where does its core advantage lie? Which industries could it potentially transform?

The key features of our sensors are the density of distributed force measurement and the deformation we can capture over the area of a fingertip. I believe we have the highest density in terms of sensing units. That is one very important metric. The other is dynamics: the frequency and bandwidth — how quickly we can detect force changes, transmit signals, and process them in real time. Other important aspects are largely engineering-related, such as reliability, drift, durability of the soft surface, and resistance to interference from magnetic, optical, or environmental factors.

A growing number of researchers and companies are recognizing the importance of tactile sensing and adopting our technology. I believe the advances in tactile sensing will elevate the entire community and industry to a higher level. One of our potential customers is deploying humanoid robots in a small convenience store, with densely packed shelves where shelf space is at a premium. The robot needs to reach into very tight spaces — tighter than books on a shelf — to pick out an object. Current two-jaw parallel grippers cannot fit into most of these spaces. Observing how humans pick up objects, you clearly need at least three slim fingers to touch and roll the object toward you and secure it. Thus, we are starting to see very specific needs where tactile sensing capabilities are essential.

From Academia to Startup

After 40 years in academia — founding the HKUST Robotics Institute, earning prestigious honors including IEEE Fellow, and serving as Editor-in-Chief of IEEE TASE — what motivated you to found DAIMON Robotics?

I have come a long way. I started learning robotics during my PhD at Carnegie Mellon, where there were truly remarkable groups working on locomotion under Marc Raibert, who founded Boston Dynamics, and on manipulation under my advisor, Matt Mason, a leader in the field. We have been working on dexterous manipulation, not only at Carnegie Mellon, but globally for many years.

However, progress has been limited for a long time, especially in building dexterous hands and making them work. Only recently have locomotion robots truly taken off, and only in the last few years have we begun to see major advancements in robot hands. There is clearly room for advancing manipulation capabilities, which would enable robots to do work like humans. While at Hong Kong University of Science and Technology, I saw increasingly greater people entering this area in the form of students and postdoctoral researchers. We wanted to jumpstart our effort by leveraging the available capital and talent resources.

Fortunately, one of my postdocs, Dr. Duan Jianghua, has a strong sense for commercial opportunities. Recognizing the rapid growth of robotics market and the unique value that our vision-based tactile sensing technology could bring, together we started DAIMON Robotics, and it has progressed well. The community has grown tremendously in China, Japan, Korea, the U.S., and Europe.

Humanoid robots assembling electronics on an automated factory production lineRobots equipped with DAIMON technology have been deployed in factory settings. The company aims to enable robots to achieve “embodied intelligence” and close the gap between what they can see and what they can feel.DAIMON Robotics

Business Model and Commercial Strategy

What is DAIMON’s current business model and strategic focus? What role does the dataset release play in your commercial strategy?

We started as a device company focused on making highly capable tactile sensors, especially for robot hands. But as technology and business developed, everyone realized it is not just about one component, rather the entire technology chain: devices, data of adequate quality and quantity, and finally the right framework to build, train, and deploy models on robots in real application environments.

Our business strategy is best described as “3D”: Devices, Data, and Deployment. We build devices for data collection, our own ecosystem, and for deploying them in our partners’ potential application domains. This enables the collection of real-world tactile-rich data and complete closed-loop validation. This will become an integral part of the 3D business model. Most startups in this space are following a similar path until eventually some may become more specialized or more tightly integrated with other companies. For now, it is mostly vertical integration.

Embodied Skills and the Convergence Moment

You’ve introduced the concept of “embodied skills” as essential for humanoid robots to move beyond having just an advanced AI “brain.” What prompted this insight? What new capabilities could embodied skills enable? After the rapid evolution of models and hardware over the past two years, has your definition or roadmap for embodied skills evolved?

We have come a long way now see a convergence point where electrical, electronic, and mechatronic hardware technologies have advanced tremendously in last two decades. Robots are now fully electric, do not require hydraulics, because hardware has evolved rapidly. Modern electronics provide tremendous bandwidth with high torques. If we can build intelligence into these systems, we can create truly humanoid robots with the ability to operate in unstructured environments, make decisions, and take actions autonomously.

“Our vision is for robots to achieve robust manipulation capabilities and evolve into reliable partners for humans.” —Prof. Michael Yu Wang, DAIMON Robotics

AI has arrived at exactly the right time. Enormous resources have been invested in AI development, especially large language models, which are now being generalized into world models that enable physical AI capabilities. We would like to see these manifested in real-world systems.

While both AI and core hardware technologies continue to evolve, the focus is much clearer now. For example, human-sized robots are preferred in a home environment. This is an exciting domain with a promise of great societal benefit if we can eventually achieve safe, reliable, and cost-effective robots.

The Road to Real-World Deployment

Today, many robots can deliver impressive demos, yet there remains a gap before they truly enter real-world applications. What could be a potential trigger for real-world deployment? Which scenarios are most likely to achieve large-scale deployment first?

I think the road toward large-scale deployment of generalist robots is still long, but we are starting to see signs of feasibility within specific domains. It is very similar to autonomous vehicles, where we are yet to see full deployment of robo-taxis, while we have already started to find mobile robots and smaller vehicles widely deployed in the hospitality industry. Virtually every major hotel in China now has a delivery robot — no arms, just a vehicle that picks up items from the hotel lobby (e.g., food deliveries). The delivery person just loads the food and selects the room number. It is up to the robot thereafter to navigate and reach the guest’s room, which includes using the elevator, to deliver the food. This is already nearly 100 percent deployed in major Chinese hotels.

Hotel and restaurant robots are viewed as a model for deploying humanoid robots in specific domains like overnight drugstores and convenience stores. I expect complete deployment in such settings within a short timeframe, followed by other applications. Overall, we can expect autonomous robots, including humanoids, to progressively penetrate specific sectors, delivering value in each and expanding into others.

Ultimately, our vision is for robots to achieve robust manipulation capabilities and evolve into reliable partners for humans. By seamlessly integrating into our homes and daily lives, they will genuinely benefit and serve humanity.

This interview has been edited for length and clarity.

❌