Normal view

Powering AI is an architecture problem

On July 22, 2026, a transmission line fault in Ashburn, Virginia—the heart of the world’s largest data center cluster—knocked more than 3 gigawatts of load off the grid in seconds. And it wasn’t the first time. Two years earlier, a single failed surge arrester dropped roughly 60 Virginia facilities and 1,500 megawatts at once. No one could anticipate so much uniform load responding to grid faults the same way, at the same time.

The AI power debate is mostly about generation: more turbines, more solar, more transmission. The grid needs more electrons. But the outages in Virginia weren’t supply failures; they were architecture failures. And a giant wave of interconnections is arriving on that same architecture, putting grid reliability at risk. It’s a problem nobody wants to own.

Asking more from the grid

The grid was built around predictable loads: steel mills, refineries, and houses at dinnertime. Different load sizes, same process—drawing power smoothly, misbehaving occasionally, and recovering gracefully.

But AI data centers don’t behave that way.

An AI campus can swing 70% of its load in milliseconds during a training run, then trip offline just as fast at the first sign of trouble upstream to protect billions in compute. Each is rational alone. Together, at gigawatt scale, they’re a problem the grid has never solved—and the next wave of data center campuses is planned at exactly that scale.

Where the old stack breaks

The standard data center power stack hasn’t changed in decades. Medium-voltage power arrives, transformers step it down, low-voltage uninterruptible power supply (UPS) units condition it, and it reaches the racks. Push that design to AI scale, and it cracks in three places.

First, the UPS sits deep inside the building, close to the racks. But its batteries are an undersized spare tire, designed to handle an outage for a few minutes, not to absorb load swings this fast and volatile around the clock.

Second, the UPS spends most of its life in bypass. Legacy converters waste enough power that operators run in eco-mode: A static switch feeds the racks directly from the grid and nothing filters in either direction. The compute’s swings go out raw, and grid transients—sub-millisecond events that can damage or take down equipment—come in too fast for any switch to catch.

Third, the protection logic was written when “large load” meant 50 megawatts. This protection logic can’t see the grid it is now a part of, so when trouble hits upstream, it does exactly the wrong thing: it drops out. In the 2024 Virginia event, most of the lost load traced to protection schemes that count voltage dips and disconnect on the third one—as designed, at the worst moment.

This isn’t sloppy engineering. It’s careful engineering the load has outgrown.

Moving into the path

The fix is three moves, made together.

Move it up—from 480 volts to medium voltage (13.8 kilovolts and higher), the voltage large sites draw from the grid.

Move it out—from the data hall to modular enclosures near the substation so the building holds only compute and the cooling that keeps it alive.

Move it into the path—instead of a battery that watches and reacts, a system every electron runs through, all the time. There’s nothing to detect and nothing to switch because nothing was ever routed around it.

On paper, three straightforward upgrades. In practice, they rewrite every line item downstream.

Making the change

When thousands of GPUs spin up together, the system absorbs the swing and hands the grid a flat load profile. When a disturbance hits, the equipment behind it never notices. A difficult neighbor becomes a predictable one. And when the utility needs help, it becomes a useful one.

Interconnection changes, too. The utility certifies one medium-voltage box instead of untangling every transformer, UPS, chiller, pump, and switchgear lineup behind it. Engineers swap chip generations without a fresh interconnection study. Months come off the permitting timeline.

Inside the fence, UPS rooms become compute or cooling space. Density per construction dollar climbs.

And the economics flip. Equipment that runs at medium voltage, sits outside, and stores its own energy can qualify for tax credits, and earn revenue in grid programs like peak shaving and demand response. Backup power stops being insurance and starts paying for itself.

The architecture test

In early 2026, we tested a full-scale system at the National Laboratory of the Rockies, a U.S. Department of Energy facility and the only place in the Western Hemisphere that can replicate real grid faults and AI-scale load swings concurrently in the same loop.

We hit it from both directions: real AI load profiles hit the compute side at full medium voltage. Grid faults hit the utility side, including a full zero-voltage event. The compute side didn’t flinch. Neither did the grid side. It cleared the large-load voltage ride-through requirements from the Electric Reliability Council of Texas (ERCOT), the grid operator, with room to spare.

Those rules exist because operators no longer take facilities this size on faith, and more are coming. Most of the industry treats them as hurdles. A medium-voltage, inline system clears them out of the box. Compliance isn’t an added feature. It’s what the architecture does.

The new layer

Much of what looks like a grid problem in the AI buildout sits inside the fence, in equipment sized for a load that no longer exists. Move the right pieces up, out, and into the path, and a grid liability becomes a grid asset. Density goes up. Permitting time comes down. Backup power earns its keep.

The engineering works—and the next wave of AI factories is being built on it. The industry hasn’t named this layer yet. We call it the medium-voltage AI UPS. The name matters less than the choice: those factories can arrive as a strain on the grid or as strength for it. We already know how to build the second kind.    

This content was produced by ON.energy. It was not written by MIT Technology Review’s editorial staff.

Architecting memory and storage in the AI era

The era of AI inference has arrived. Imagine a healthcare system analyzing millions of data points in real time to accelerate life-saving medical research, or an intelligent assistant instantly resolving thousands of complex customer needs at once. These real-world breakthroughs rely on advanced infrastructure acting as the engine of continuous intelligence, powering real-time services while also supporting an increasingly intelligent edge of IoT and consumer devices. However, in this inference-driven landscape, every delay, bottleneck, or wasted watt directly affects human outcomes and operating costs. 

This shift changes what infrastructure must deliver. Performance, latency, memory bandwidth, storage throughput, and networking cannot be optimized in silos. Inference workloads are continuous, geographically distributed, and highly sensitive to response time, requiring systems designed for scale, resilience, and efficiency from the start.

“We tend to think of AI as a single workload, and it’s not. It’s thousands, it’s millions, it’s billions of different workloads,” says Jim McGregor, founder and principal analyst, Tirias Research. AI inference changes the optimization problem from one of raw compute to coordinated infrastructure—memory, storage, and networking.

For business leaders, the priority is clear: AI infrastructure decisions must balance cost, flexibility, and future readiness. The winners will be organizations that improve performance per watt, reduce environmental footprint, and remove memory and storage bottlenecks before they limit growth.

AI inference requires a new architectural approach

Systems for AI need to be rearchitected because shoehorning modern AI systems into legacy infrastructure limits AI’s transformative potential. Purpose-built architectures are essential to realize the true value of AI, from accelerating scientific discovery to creating truly autonomous digital agents.

Traditional enterprise IT has been able to rely on relatively stable infrastructure assumptions, but inference and agentic AI introduce new demands around latency, data movement, scalability, and utilization that make architecture choices far more consequential.

“Data centers must now support continuous, distributed, and increasingly real-time AI services—none of which are a single workload,” says McGregor. “They all require different requirements from a system-level perspective.”

To support real-time AI, enterprises can no longer view memory and storage merely as supporting hardware, but at the heart of the system. Organizations need to architect a data pipeline that can rapidly ingest, clean, transform, store, move, and deliver data. Inference workloads place sustained pressure on infrastructure in ways that look very different from earlier training-centric deployments, demanding continuous data retrieval and caching that traditional applications never required.

Accordingly, performance by itself is no longer the sole benchmark that matters. Enterprises increasingly must balance performance with efficiency, cost, and scalability, especially as they try to support different AI services without overbuilding infrastructure for peak conditions.

“You have to optimize the entire network, and that includes memory and storage, around the types of workloads you plan on running,” says McGregor. “You have to really have a detailed understanding of what those workloads are going to be.”

Any AI infrastructure strategy must start with workload awareness. Inference, agentic AI, and other emerging AI use cases require organizations to treat the data center as an integrated system.

Data movement is the new bottleneck and an opportunity for competitive advantage

As enterprises deploy advanced inference and agentic systems, the sheer volume of data being queried in real time has made data movement the most pressing constraint. Modern AI techniques like retrieval-augmented generation (RAG) require systems to constantly scan massive databases to generate accurate responses. This requires immense computing power, but more importantly, it requires immediate access to data.

McGregor says the focus shift to how efficiently data can be moved, cached, and delivered across the broader architecture elevates memory and storage from background infrastructure to strategic assets. “The biggest thing we’re doing right now is moving data from one place to another and making sure that we can use it effectively.”

Because AI is not a single workload category, simply buying the fastest processors is insufficient. Inference depends heavily on memory bandwidth, caching, storage proximity, and the ability to retrieve relevant information quickly and consistently. Understanding where each resource belongs in the stack and how those layers interact under real operating conditions has become a business imperative.

The most effective AI infrastructure looks less like a collection of best-in-class parts and more like a balanced system of compute, memory, storage, and networking, McGregor says, because bottlenecks tend to migrate from one layer to the next. “You have to architect all four together to be efficient, and that’s the challenge.”

The interdependence of data-plane design and network bandwidth means AI infrastructure planning has become a business decision just as much as an engineering one: latency is now inseparable from value. In robotics, financial services, healthcare, and customer-facing AI systems, delays are not merely technical imperfections; they can undermine safety, responsiveness, or trust. AI infrastructure performance becomes a matter of reputation management.

The organizations that gain the most from AI may not be those with the largest clusters, but those with the clearest understanding of how to align every infrastructure element to effectively execute AI workloads.

Building an AI infrastructure procurement framework

Planning AI infrastructure is not simply about choosing the fastest hardware. It is about how to scale without locking the organization into assumptions that may quickly become obsolete. “You need to be flexible because the demands are going to change rapidly and the technology is changing rapidly,” McGregor says.

Future-proofing AI infrastructure requires keeping your options open as workloads, economics, and architectures keep shifting:

  • Define the AI workloads that are being optimized. Infrastructure choices must match business needs rather than what McGregor calls generic “AI readiness,” which risks overspending in some areas while leaving bottlenecks unresolved in others.
  • Build a modular architecture for compute, memory, storage, power, and cooling so capacity can change as demand shifts rather than committing too early to a rigid architecture.
  • Work with the full ecosystem of suppliers and integrators to reduce supply risk and improve access to the right components. McGregor says buyers can no longer assume their OEM or cloud provider alone will insulate them from supply constraints or architectural complexity.
  • Reassess your procurement strategy continuously. AI requirements, hardware, and business models are changing too quickly for a fixed long-term design.
  • Optimize for efficiency and ROI, not just peak performance. The most powerful setup may be too costly to sustain. Efficiency is also a public-facing metric—better utilization and more workload-aware system design can help companies respond to growing scrutiny around power consumption and water use.

The strategic goal of smarter AI data center design is not maximum performance at any cost, but an adaptable architecture that can deliver value, absorb change, and justify its footprint.

AI infrastructure is now a business strategy

AI data centers have quickly evolved from a back-end technical concern to becoming strategic business systems that help determine how effectively an organization can turn AI into revenue, improve human outcomes, and create a competitive advantage.

In the inference era, memory and storage are no longer passive repositories, explains McGregor, they are the active lifeblood of AI. The organizations that gain the most from AI will not necessarily be those with the largest computing footprint, but those that align infrastructure investments to business outcomes, reduce data bottlenecks, and build the flexibility to adapt as workloads evolve. He predicts that competitive advantage will increasingly belong to enterprises that treat compute, memory, storage, and networking as an integrated system designed to deliver AI efficiently, at scale, and with measurable ROI.

Procurement is now strategy and system design is a leadership issue, McGregor concludes. “One of the biggest questions every executive has to ask is how is AI going to change my business model?”

This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight.

Facilitating AI integration with simplicity at scale

As companies scale, the technology supporting operations can become a liability just as quickly as it becomes an asset. Disconnected systems, site-specific tools, spreadsheets, and manual workarounds can create data silos that make it harder to spot problems early, coordinate responses, and make decisions with confidence. For Jabil, a global manufacturing company with more than 100 sites across more than 30 countries, the answer has been to make integration and simplification a priority.

The company adopted a “simplify-first, then-innovate mindset,” says Harish Manohar, SAP IT director at Jabil, recognizing that adding new technologies without first reducing complexity risks creating more risk. The goal is to standardize processes, consolidate where possible, and establish a more consistent data backbone across the organization. “Any innovation without simplification is going to add more complexity,” Manohar says.

That philosophy also changes how Jabil approaches modernization. “Any modernization or transformation should add measurable business value,” Manohar says. The company is focused on connecting processes end-to-end across its supply chain and creating a foundation that can scale consistently across regions. Integration comes first because, as Manohar puts it, “the backbone of any contemporary or modern organization is data.” Before organizations can optimize, automate, or apply AI, data needs to flow seamlessly across systems.

But doing that across a global organization is hardly straightforward. Jabil’s more than 100 sites operate with different levels of process maturity, legacy systems, and localized workflows, while regulated businesses bring additional compliance requirements. As such, standardizing across different regions and business environments means changing processes and governance without disrupting the operations already in place.

The value of that work extends beyond the technology to the people using it. Integrated workflows can offer employees shared visibility into data, reduce manual data reconciliation, and help them move from chasing information to acting on insights. For Jabil, the aim is also to improve real-time visibility into supply chain events, which can enable faster responses to disruptions and reduce operational risk.

Looking to the future, that foundation could make AI and automation all the more useful and scalable. With trusted data and integrated systems in place, Jabil is exploring predictive supply chain insights, intelligent exception handling, and AI-driven planning and forecasting. To Manohar, the takeaway is clear: “Simplicity at scale is a very competitive advantage,” and technology investments must ultimately connect to business value and operational resilience.

This episode of Business Lab is produced in partnership with SAP.

Full Transcript:

Megan Tatum: From MIT Technology Review, I’m Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace.

Our topic today is enterprise technology integration, and how the benefits of consolidating tools and systems across the supply chain help organizations operate more reliably at scale. When companies reduce tool sprawl and connect their systems more effectively, they gain earlier visibility, faster response, and greater resilience across production lines.

My guest today is Harish Manohar, SAP IT Director at Jabil. Jabil has been on a journey to simplify its technology landscape by using SAP Integration Suite as the foundation to connect systems, retire fragmented tools, and enable more consistent operations globally.

This podcast is produced in partnership with SAP.

Welcome, Harish.


Harish Manohar: Hello, Megan. Good morning.

Megan: Thank you so much for being here, Harish. Just to start, if we could set some context, can you give us a quick overview of Jabil, the business and its overall transformation journey?

Harish: All right. So, about Jabil. Jabil is a global manufacturing company headquartered in St. Petersburg, Florida, USA. We have about 60 years of experience offering comprehensive engineering, supply chain, and manufacturing solutions across different industries. We have a global footprint of about over 100 different sites across 30-plus countries, 140,000-plus employees.

We are a trusted partner for more than 400 of the world’s top brands. That’s a little bit about Jabil.

Megan: Fantastic. And a lot of scale there, as you’re referring to some of the stats there. As things have got more complex, where did disconnected tools and systems start to slow you down, and what ultimately drove you to make integration a really strategic priority?

Harish: I talked about our global footprint across 100-plus sites. With a global footprint always comes complexity about site-specific tools. All of our sites have been in business for a long time, and over the period of years, they had their own tools for their own processes. It’s a little bit disconnected.

When we are looking to scale, the first thing we wanted to start looking at is what is this mix of site-specific tools, manual workarounds, spreadsheet-based processes, legacy applications, whatnot? That’s a big technical debt that we have had over the last 25 years. That’s where we started, and that is what led Jabil to make integration a strategic priority because we had limited ability to see issues early across plants, across regions, which would help us to coordinate responses consistently, and also to be able to scale those responses consistently. This complexity created data silos, and in a way delayed robust decision-making.

As the complexity increased globally, integration became extremely critical to a few things. It was critical to establishing a single trusted data backbone, which would directly enable faster coordinated responses across the network. We at Jabil, as part of our transformation journey, believe that having the right data at the right time fundamentally changes how we respond to disruptions, which is all about a manufacturing business. How we respond to disruptions. This is where we started our strategic priority towards having a integrated system that drove data consistency across our landscape.

Megan: Fantastic. As you have outlined there, there was obviously a real commercial need for this, but how did you think about bringing in those new technologies without adding even more complexity to the mix?

Harish: Great question. Whenever we talk about transformation, we talk about all these bleeding-edge technologies that are out there today when it comes to AI and data, cloud, et cetera. But it was very important for us to put a stake in the ground and say and adopt a simplify-first, then-innovate mindset. Because any innovation without simplification is going to add more complexity, just like you mentioned.

For us, SAP is our core digital platform. We want to focus on bringing more processes into SAP as much as possible. That is easier said than done because we have been in business for a while, global company, so not all processes exist within SAP at this point in time. We are slowly trying to standardize those processes, and having them under one single source of data would help us scale faster in terms of having data silos. We don’t want data silos across different systems.

This is where we started to introduce newer capabilities around SAP. From a cloud standpoint, we have been using SAP’s BTP and Integration Suite, which is proving to be the center stage of all integrations across Jabil. Well, it’s not there yet, but that is the direction that we want to pursue is we don’t want to have a slew of different integration platforms, rather try and see where Integration Suite fits best and where other smaller integration platforms would add more value.

Similarly, we have adopted an API-driven, event-based integration approach. That is our best practice that we have put down because we don’t want to keep moving data from one place to the other. That’s not good business practice in the IT world. Most of our integration architectures are API-driven and event-based. That is our focus.

Coming back to your new technologies perspective, we want to reuse as much as possible and standardize versus going out and buying these one-off tools that solve for point-in-case use cases. We really don’t want to go down that path. For major processes, we do adopt a best-of-breed approach, but for, let’s say, site-based use cases where a specific site has a particular need for a tool, we try and standardize that and reuse what exists in a different site, for example. There may be some need for a business process change, minor process changes, but that is our direction to make those process changes and reuse what is there already in a different location or a different region. So, that’s one.

Lastly, we are heavily aligned with SAP’s clean-core approach when it comes to customization. That has been the challenge for us over the last 25 years where we have been using SAP is our systems are heavily customized because we cater to different customers across the globe. Most of our demands are customer-driven, so we have to put in play these heavy customizations.

But now we are taking a pause, and we are saying, “You know what? We have customized so much so far, but now we are moving our systems into RISE, which would enable a clean-core journey in the future.” Now we have to put really good governance criteria and review processes that do not allow heavy customization of our system. We want to move away from that model as much as possible. Again, it’s not easy to do that at this point in time, but there is always a start.

Megan: I mean, it sounds like you took a very incremental, intentional approach to this. I mean, as you scaled globally then, what did modernization really look like at the company, and why start with integration?

Harish: To that point, we have always looked at transformation, modernization, very objectively. For us, it’s just not about upgrading a system. That’s not what it is. Any modernization or transformation should add measurable business value is our model, is our charter. Having said that, we don’t look at modernization in terms of just upgrades, but in what it gets our business in terms of value.

Most of our modernization transformation approaches are focused on connecting processes end-to-end across our supply chain, which is key for our business value. And then we also have a very concerted effort going on in the business community: how to standardize how our plants operate globally. Because, like I mentioned earlier, we have 100-plus plants, different processes, different legal regulations, different countries. It’s very hard for us to come up with one template across the globe, but we are trying to standardize as much as possible. And that’s where we are leveraging SAP’s Signavio, which is our business process management tool. We want to leverage Signavio’s capabilities in helping us standardize these global processes.

Now, back to your question, why did integration come first? Because the backbone of any contemporary or modern organization is data. And to get the right data at the right time, integration is the key aspect of the whole optimization exercise. Data needed to flow seamlessly before we start optimizing or automating or even applying AI use cases. This is where integration came first.

We wanted to create a single system of record across the operations. Well, when I say “single system of record,” it’s not just SAP, but the ability for us to create those data pipelines across those systems of record being supply chain, planning, inventory, et cetera, et cetera, in that operation space.

The result is we want to get to a foundation that helps us scale consistently across our different region. That is our main objective is to, how do we scale as the business grows, as we develop into this bigger organization across different industries? How do we set this foundation that will help us scale consistently? Simplicity, consolidation becomes strategic assets at scale.

Megan: Absolutely. And you touched on some of the complexities there of doing this at a scale that Jabil is at with its international footprint. What were some of the biggest challenges in your view in terms of rolling this out across regions, and how did that more standardized approach that you’ve mentioned there help?

Harish: Absolutely. I would like to reiterate some of those key challenges I mentioned. One hundred-plus sites, different sites have different maturity levels in terms of how they approach processes. They have a multitude of different legacy systems, localized processes, workarounds, spreadsheets, and the change management that exists within each site is very different. And we do have a footprint of highly regulated businesses. And when it comes to regulated businesses, that comes with its own set of challenges around qualification and CSD processes, et cetera. These are the key challenges that we are up against.

Now, the standardization helped us provide consistent workflows, data flows, and governance across sites. Now, we are not there at 100%, but we are working towards that, providing consistent workflows, data pipelines, and governance across sites. And we want to enable faster rollouts of our new bleeding edge technologies. For example, when I talked about SAP’s BTP or SAP Signavio or any other new SAP tool or non-SAP tool, traditionally our ability to deploy those had a challenge around the heavy customization that is required for each and every site. Now, the standardization approach takes that heavy customization out, which enables a faster rollout of those newer technologies.

And then lastly, we want to scale across all of our plants. I think initially we want a target of about 40-plus plants with shared processes that are consistent across the different regions. We want to shift from a site-by-site operations model to more of an enterprise-capability approach.

Megan: Right, and fascinating. And you touched on the people management aspect of this as well, because obviously this isn’t just about technology, it’s about people too. So, from the employee side, how did this shift to a more simplified landscape change the day-to-day experience for people compared to juggling multiple tools at once?

Harish: Great question. And I’ve been hearing direct feedback from our business community on some of these transformation initiatives on how those have changed their daily jobs significantly. Before we embarked on this transformation journey, any employee, any persona. You take a buyer, you take an inventory planner, you take a finance analyst, we go by personas. They had to deal with multiple tools, manual coordination, data reconciliation, especially in the finance space, inconsistent processes across different regions. And then the time spent reconciling data resulted in delay of making decisions, robust decisions. This was the before.

But now, since we are moving towards this newer standardization and more of an integration approach, we are able to achieve, to a certain extent, a single integrated workflow across different systems. We have built some key processes that will enable the single integrated workflows across systems. The users, our business community, irrespective of their roles in the organization, have clear visibility and shared data across their teams, which is very important. Earlier, they were dealing with different versions of the data, local workbooks, spreadsheets, and then the time spent talking to each other and reconciling what is the right data? What is the single source of truth? That we are trying to peel away that layer and get to that where the employees don’t have to deal with that kind of complexity.

This reduces manual effort and enables faster issue resolution when it comes to actual disruptions. Employees move from chasing information to focusing on acting on insights, which is where I think the new age of AI comes into play. I’ll talk about that in a little bit, but technology becomes an enabler of decision-making, and it’s no more an overhead. That’s where we want to go.

Megan: Fantastic. Such an important element of this, isn’t it, that people side of things? And we’ve touched briefly on this idea of value you’ve talked about before, because with an initiative of this scale, ROI is always front and center, of course. What benefits stood out most for you, and how important was better visibility in particular across systems?

Harish: Megan, I talked about how modernization and transformation for Jabil means measurable business value, which is directly connected to the ROI. We don’t do any transformation initiatives just because we want to do it from an IT standpoint. Any investment that we make in a transformation or a modernization initiative has to have a deliverable business case that is approved, signed off by business, because that is the only way true transformation happens, if IT and business are a partner as part of this transformation journey.

The biggest benefit that we have seen in this initiative is we are striving to reach, attain real-time visibility across our supply chain events. That is the biggest benefit that we see, faster response to disruptions and exceptions. And we are working to reduce our operational risk significantly by operating in this newer model.

One example I can give you is the unified workflows that I talked about earlier. It enabled earlier identification of missing materials and faster resolution across our sites. When it comes to a manufacturing company that has a global footprint, materials are the backbone of our whole supply chain process, right? Having a unified workflow, which is able to identify missing materials early in the game, was a game changer for our whole operations community.

Real-time analytics allow instant supply chain adjustments without delays. We are focusing a lot on getting analytics, a global analytic footprint in place that allows instant supply chain adjustments without any delays. That’s where AI is going to play a major role currently, and also in the near future.

And again, when you talk about visibility. Visibility is not just about reporting what is there in the system. Visibility directly should enable scenario modeling for our users to make strategic adjustments in their processes, which visibility also should make way for proactive decision-making, and also foster business continuity. This is how we look at visibility at Jabil.

Megan: Right. And you’re still on this journey, of course, but now that Jabil has a strong integration foundation in place, what does it unlock next for you, and how are you thinking about AI and automation as you’ve touched on a couple of times?

Harish: Yeah, we have talked about a couple of times around AI. So, we strongly believe at Jabil, a strong integration foundation enables event-driven, real-time processes, robust decision-making, scalable automation, and all of this enable easier adoption of AI use cases. And again, we are in the new age of AI. We are working towards getting to a AI-enabled enterprise, but having these foundations in place truly fast tracks our approach of AI use cases.

Our key focus areas, when I’m thinking about AI and automation in the immediate future, are predictive supply chain insights, intelligent exception handling, which is key to our business operations from a site operation standpoint. Intelligent exception handling is very, very key. On the supply chain side, I talked about predictive insights. That is also absolutely important. All of this enables AI-driven planning and forecasting capabilities.

For us, AI should augment true decision-making and robust decision-making, and deliver measurable value, not just experiment AI in use cases. We want to move beyond just experimenting AI in our business processes, but we want that AI that we implement to truly augment the decision-making process that we have, and also deliver key business value.

How does all this connect to integration? Integration ensures AI has access to trusted data, and also enables the ability to act across multiple systems in a global company like Jabil.

Megan: Fantastic. And if we could just finish, I suppose, with a little bit of advice for others, for other leaders, perhaps, dealing with tool sprawl at the moment, what are some key lessons you would say you’ve learned about prioritizing integration right from the start?

Harish: Absolutely. When it comes to tool sprawl, we can go all day about what are the different areas of tool sprawl? For example, application development, we have a multitude of tools; via integration, we have a multitude of tools. Data, we have a multitude of tools, but let’s just focus on integration. That’s the core topic here.

I would recommend folks that are in transformative roles in their organizations to start with integration as a foundation and not as an afterthought, right? Prioritize simplification over adding a slew of different tools to address different capabilities. Try and simplify as much as possible before we start your upgrade or your transformation journey. Standardize over locally optimizing tools. Try and get to that. Try and get the business community, your key SMEs in the business space, to understand the value of standardization and simplification of processes and how that enables your business to deliver value faster.

Second one, after prioritize: build a single source of truth for data as much as possible. I’m not saying it’s going to be always the case where an organization as in the scale of Jabil will be able to function just with SAP. They’re going to have different systems, but try and get to a model where you’re working with a single source of truth and not locally siloed data sources, right?

Next is focus on building a scalable integration architecture. Don’t just confine yourselves to the current state where you are, and build something in place that will only serve you for the next six months to a year. No, that’s not the goal. Anything that you build as an integration architecture should be scalable, and should serve the organization for the next three to five years. That’s how I look at it. When I’m putting in a new architecture pattern or a new event-driven insights, I look at, “Okay, where is Jabil going to be two years, three years down the line? Would this suffice for that scale?” That’s how I look at it.

Then focus on outcomes and not just technology. Focus on outcomes: speed, visibility, resilience, and not just technology deployment, because end of the day, IT and business should partner on the business value and not just technology upgrades.

Going back, simplicity at scale is a very competitive advantage, and technology investments must tie directly to business value and operational resilience. That’s how we look at Jabil in terms of our tool sprawl and how we prioritize integration right from the start. And that’s what I would suggest to other leaders that are looking to advance in this space.

Megan: Fantastic. Brilliant and very comprehensive advice. Thank you ever so much, Harish. And thank you ever so much for joining us. That was Harish Manohar, SAP IT Director at Jabil, whom I spoke with from Brighton in England.

That’s it for this episode of Business Lab. I’m your host, Megan Tatum. I’m a contributing editor at Insights, the custom publishing division of MIT Technology Review. We were founded in 1899 at the Massachusetts Institute of Technology, and you can find us in print, on the web, and at events each year around the world. For more information about us and the show, please check out our website at technologyreview.com.

This show is available wherever you get your podcasts. And if you enjoyed this episode, we hope you’ll take a moment to rate and review us. Business Lab is a production of MIT Technology Review, and this episode was produced by Giro Studios. Thanks so much for listening. Goodbye.

This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight.

Unlocking hidden revenue streams with market models

Each day, an airline transports tens of thousands of passengers on hundreds of flights. Often these are not straightforward point-to-point routes, with passengers requiring multiple connections. The airline can consider potentially hundreds of variables to price each of these journeys: demand, season, time of day, current events, global markets, and competitor airline activity to name just a few. It is a nuanced process that must constantly adapt to the goings on in the wider world.

Generative AI-powered market models are emerging as a means of handling complex tasks like this in real time. These deep learning models are trained on high-resolution numerical data and designed to analyze, simulate, and predict complex financial dynamics. Rather than relying on historical trends or static rules, the market model acts as an AI “brain,” consolidating a variety of data to simulate different market environments and make dynamic commercial decisions, such as pricing, inventory, or revenue management.

“It helps us make better, faster, more granular commercial decisions,” says Dominic Kennedy, senior vice president of revenue management, sales, and e-commerce at Virgin Atlantic about the market model his team is using to drive their generative pricing engines in some markets.

“It considers, on a real-time basis, a plethora of different inputs, whether it be demand, capacity, or booking. It has a really sophisticated way of evaluating our positioning relative to competitors, market conditions, and a whole raft of other things that have significance in how demand is manifested,” he adds.

Download the full report

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Agentic AI has a latency problem that more compute won’t solve

Abstract neon pattern of distorted purple columns and lime-green rings rippling across a dark background.

Half of enterprise AI deployments are missing their own latency targets at peak load. This is the headline finding of Akamai’s State of AI Inference 2026 report, which surveyed 200 AI practitioners and found that 82% of organizations say their most critical use cases require end-to-end response times of 500 milliseconds or less. A total of 64% of organizations now require end-to-end response times of less than 250 milliseconds for their most important use cases, yet 50% of deployments are failing to meet these latency demands at peak load.

My colleague Ari Weil, who leads product marketing for our cloud computing business and ran point on that research, summarizes the findings well: “The enterprise AI honeymoon phase is over… they are hitting the latency wall.” 

Agentic workflows aren’t a “single round trip”

The latency issue stems from the way agents work. It’s an iterative process, somewhat like a king sending out knights, emissaries, and messengers to conduct the business of the kingdom. There are many comings and goings, not just one person sent on a single round trip. 

For instance, when an agent built on a framework like LangChain, CrewAI, or Pydantic AI received a user request, it can fan out into dozens of sequential operations such as a reasoning call, a tool invocation, an API lookup, or a context retrieval. Then an agent may execute another reasoning call to decide what to do with what just came back. Every one of these operations or “hops” that must cross a wide-area network to reach a centralized data center adds latency, and a chain of 50 hops can multiply that transport time into seconds on its own, regardless of how fast the model generates tokens.

In fact, in a paper posted to arXiv in November 2025, researchers found that CPU-side processing can account for up to 90.6% of total latency in agentic workloads. In other words, your GPU might finish a reasoning step in a few hundred milliseconds, but then it might have to wait on additional tool call runs to CPUs in distant data centers. This is what causes spikes in GPU idle time. 

“More GPU capacity does nothing for this. You can’t brute-force your way out of a wait state.”

More GPU capacity does nothing for this. You can’t brute-force your way out of a wait state. This is the part of the conversation the industry keeps skipping, mostly because “buy more GPUs” is a much quicker fix to suggest than “figure out where your CPU-bound work is actually executing and why it’s so far from the data it needs.”

We need new benchmarks to fix the latency issue

One reason the looming latency wall sneaks up on teams is that they are not looking at the right benchmarks for agentic workloads. Most LLM-serving benchmarks measure tokens per second and GPU utilization on a single box. That’s great if the workload is indeed on a single box (i.e., one model answering one prompt), but that’s not the case with agentic workloads. Those benchmarks don’t address an agentic response that, say, makes a 50-hop chain cross a WAN 4 times to reach 4 separate services. 

“Staging may pass the benchmarks because it tests the model, but production tests the whole chain, including every hop your serving engine was never designed to see.”

That’s where the gap lies: Staging may pass the benchmarks because it tests the model, but production tests the whole chain, including every hop your serving engine was never designed to see.

The 500ms wall is not a soft target

This is showing up at scale because agents are moving into production faster than most teams’ architecture is evolving to support them. LangChain’s State of Agent Engineering 2026 survey of more than 1,300 professionals found that 57.3% of organizations now have agents running in production, up from 51% a year earlier. Among those builders, latency has become the second-most-cited barrier to production, behind only output quality. 

This is a serious issue for application teams. The 500ms threshold in Akamai’s survey isn’t a performance goal teams can afford to miss. For a live customer interaction or a real-time compliance check, that 500ms determines whether the application works or it doesn’t. 

We’ve solved this problem before

There’s a reason this feels familiar to anyone who was building for the web in 1999. Akamai exists because of a nearly identical problem. MIT researchers Tom Leighton and Danny Lewin founded the company to answer a challenge posed by Tim Berners-Lee: fix what the press had started calling the “World Wide Wait,” the crushing latency of pulling every request back to a small number of centralized servers. When the trailer for The Phantom Menace crashed sites across the internet in 1999, the culprit was distance: millions of browsers all reaching for the same far-away origin server at the same moment. The fix moved content to thousands of points closer to the people requesting it, instead of trying to build a faster origin.

Agentic AI is running into the same wall, just in a different vehicle. AI works just fine on centralized inference if you’re talking about running batch jobs overnight. But today’s applications built on agentic AI are real-time loops sitting inside live transactions, and the fix for agentic lag is distribution. Instead of expanding racks of CPUs and GPUs at the center, we need to move agentic execution to where the model’s tools, context data, and users actually live.

Agentic AI needs a tiered architecture, not a bigger data center

In practice, agentic AI requires a tiered architecture, one that includes a centralized core, regional GPU clusters, and CPUs at the Edge. 

  • Centralized core—perfect for heavy reasoning over large context windows, where the round trip to a large model matters less than the model’s raw capability.
  • Regional GPU clusters, increasingly built on hardware like NVIDIA’s Blackwell platform—ideal for localized inference, so the heaviest compute sits closer to where demand actually concentrates.
  • Edge CPUs—the essential component for speed. This is the nexus for tool execution, orchestration, and context retrieval, since these are the steps that happen most often in a chain and benefit most from sitting next to the data and APIs they call.

We’ve built Akamai Inference Cloud around this tiered framework. It’s the same distribution logic behind our AI Grid Orchestrator. We route CPU-bound orchestration and tool calling to the edge, and keep GPU-bound reasoning where it makes sense, regionally or centrally. 

What to demand before you commit

The good news is you don’t need to distribute every workload to the edge on day one. But before you commit to a production architecture, you should know which of your agent’s dozens of hops are latency-sensitive and which aren’t. Then build a defined performance budget for each one.

“The teams that treat it as a GPU-shopping decision will be back here in six months, staring at the same four-second response time, wondering why more compute didn’t help.”

My advice is this: Before signing off on a large-scale inference deployment, ask your infrastructure for four things: 

  1. Portability across regions and providers
  2. Elasticity to absorb peak load without falling over
  3. Data locality so tool calls aren’t crossing oceans to reach the context they need
  4. A performance budget you’ve actually tested against production traffic, not staging traffic.

The teams that address this infrastructure decision now will be the ones whose agents still work when the benchmark environment transitions to real users. The teams that treat it as a GPU-shopping decision will be back here in six months, staring at the same four-second response time, wondering why more compute didn’t help.

The post Agentic AI has a latency problem that more compute won’t solve appeared first on The New Stack.

Scaling AI agents with trustworthy data

Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potential to transform work. But many organizations find that realizing the desired return on investment (ROI) from AI hinges on having the right foundation, with inadequate infrastructure and data being major blockers.

Agentic AI places considerable new demands on enterprise data systems. The shift from answering questions to taking actions means AI agents need data from across the enterprise, in all its structured and unstructured forms, and with the right business context. To make decisions and act in real time, agents also need frictionless access to the organization’s operational systems—for example, those storing its supply chain, point-of-sale, or human resources data. Legacy data systems, even those updated just a few years ago, struggle to meet these demands.

As AI agents become embedded more widely in enterprise operations, the need to overcome the restrictions of legacy data systems grows more urgent. If Gartner’s prediction that AI agents will augment or automate 50% of business decisions by 2027 proves correct, organizations must eliminate bottlenecks or risk depriving agents of the data they need to make the right decisions at speed.

This report, based on a survey of 300 data and technology executives, explores how legacy systems are limiting the effectiveness of AI agents in many organizations. It finds that a handful of organizations—the data leaders—are having greater success with agentic AI and experiencing fewer data limitations as a result of legacy systems. These leaders offer a guide to creating the right data environment for agents to flourish and trusted systems to scale.

Key findings from the report include:

Few companies currently provide agentic AI with ample access to enterprise data. Across all the surveyed organizations, AI only has access to an average of 45% of company data. That number falls to 30% or less in organizations categorized as “data laggards”. A select group, however, ensures access to over 70% of their data. These “data leaders” are having greater success with their agents than the rest.

Trust in agent decisions is a reflection of data readiness. Today, only around half of surveyed organizations trust that the decisions their AI agents make are accurate and relevant. By contrast, 100% of the data leaders trust their agents’ decisions, a strong indicator that reliable AI requires a reliable data foundation.

Data leaders find it easier to achieve agent scale and speed. Two-thirds of data laggards say legacy data systems limit AI agent scaling (66%) and prevent agents from making decisions at speed (68%). Having largely overcome legacy data constraints, the leaders have mostly cleared these roadblocks, with just 8% reporting either constraint.

The pressure is on to make data estates agent-ready. Within two years, 100% of respondents plan to be using agentic AI, with 69% expecting to use it widely. Without removing data system constraints, agentic AI will fail to deliver the desired speed and efficiencies it promises.

Data access and context are top priorities. The most important initiative to enable scaling among all respondents is improving access to structured and unstructured data for AI agents. Also high on the list is enhancing data and AI governance with business context. Data leaders are also focusing heavily on the automation of data management.

Download the full report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

“Just rewrite it”: What platform teams really think about modernization

Colorful illustration of a diverse crowd of people with varied hairstyles, clothing and expressions gathered closely together.

Mergers, acquisitions, and the steady churn of business and technology initiatives are creating something nobody asked for: Duplicate infrastructure and expertise. 

Here’s the typical split: A platform engineering team that owns cloud-native and Kubernetes workloads. Meanwhile, traditional IT holds the keys to virtual machine (VM) workloads. Two teams. Two domains. One budget. And the costs keep going up.

Even organizations that talk about standardizing on Kubernetes still have a substantial VM footprint. For many teams, this coexistence isn’t a temporary transition state. It’s the operating model.

On-premises, this split forces two separate environments. Each environment includes networking, servers, and storage. Such duplication can be structurally less cost-efficient than consolidation. VM-based mission-critical workloads aren’t going away anytime soon.

It’s not like teams don’t want to modernize. They absolutely do. But it’s not as simple as just picking between old-school VMs or diving into Kubernetes. What’s really happened is these two worlds have grown up on their own.

That kind of split often leads to extra infrastructure, more people doing the same jobs, slower projects, mixed-up governance, and budgets that keep ballooning. And when you’re on-prem or working at the edge, running two separate setups for networking, compute, storage, and playbooks just doesn’t make sense anymore.

To make matters worse, “just rewrite” bares its fangs on the modernization initiative. Finance and executive leadership see two teams running two tech stacks. It’s only natural that they reach for the obvious fix: Pick one team’s platform, with no technology consideration, migrate everything to it, and watch the added cost disappear from the executive briefing slide and move to the CFO’s budget spreadsheet.

Rewrites are rarely the shortest path to business value

Over time, we learned from our customers that “rewrite it” isn’t a modernization strategy. Rather, it’s a budget, risk, and timeline strategy all at once. In many cases, rewrite it doesn’t make sense financially. The tech industry loves the idea of re-platforming and re-architecting legacy applications. Even then, such a move only returns your enterprise to square one and functional parity. The more realistic path is to keep mission-critical applications as-is when scaling out cloud-native platforms to deliver new value.

Moving everything to Kubernetes/cloud initiatives won’t prevent two platforms either. Such initiatives often stall because some workloads don’t fit or take far longer than planned.

The economics of rewrites don’t disappear just because AI accelerates software delivery. AI can compress the time it takes to write code. However, writing code was never the expensive part of a rewrite. The costs that dominate many rewrite budgets are judgment costs, and those remain stubbornly human.

Start with architecture. Organizations still need software engineering expertise to design the target system. That design problem has gotten harder, not easier. Cloud-native applications built on microservices for horizontal scaling bear little structural resemblance to the traditional enterprise applications they replace. Someone has to make those translation decisions and then spend the time directing the AI on what to build. That direction time is a real line item.

Validation is the next cost that survives. When customers or employees depend on a piece of software, even small behavioral changes are disruptive, making it non-negotiable to prove feature parity. Testing and validating that parity remains heavily human work. AI can generate test cases. It can’t tell you which broken workflow will cost you a customer.

Then comes the data. Teams must migrate and adapt data to the new system, and that work almost always surfaces complexities nobody scoped, including undocumented dependencies and format assumptions baked into decades of records. No amount of generation speed on the code side makes the data side move faster.

The rewrite math changes shape with AI. It doesn’t shrink to zero. The spend shifts from writing software to decision-making, verification, and migration.

The rewrite math changes shape with AI. It doesn’t shrink to zero. The spend shifts from writing software to decision-making, verification, and migration.

We see the same pattern repeat with rewrites among our customers. They keep mission-critical systems running as they are. Then they build new value with cloud-native applications in parallel. Modernizing selectively only when it’s truly worth it.

The real gap is operational 

The gap we see isn’t philosophical — VMs versus containers — it’s operational. The tooling, workflows, and skills that define VM and cloud-native operations differ. If platform teams can’t deliver these services at the expected velocity, developers will blame the platform. When developers are accustomed to provisioning core services in minutes, any friction in on-prem or edge environments is perceived as the platform adding friction or slowing delivery.

The gap we see isn’t philosophical — VMs versus containers — it’s operational.

Historically, day-to-day operations in VM environments are UI-driven. Cloud-native environments are much more command-line interface (CLI) driven, where APIs, config files, and the terminal are the center of gravity. That gap becomes both an organizational and technical constraint. Moving from UI-driven operations to deep command-line interface (CLI)/config workflows isn’t a natural step without a significant shift in the team’s capabilities.

The operational gap shows up quickly in data services. Cloud-native workloads don’t just need compute. They need databases, object storage, file, and block services delivered at cloud-like speed. And despite the myth that containers are stateless, the reality is that most meaningful workloads have state somewhere as data, logs, metrics, or dependencies that must be handled consistently.

Another notable gap is that storage consumption differs: 

  • Cloud-native apps often need multiple storage types simultaneously
  • VM workloads historically rely on straightforward block storage

The public cloud, by shaping cloud-native expectations, further contributes to the gap. Developers can click to get a database, such as Amazon Relational Database Service (RDS), and object storage, such as Amazon Simple Storage Service (S3), is just there. Developers expect this level of self-service simplicity when these platforms are extended beyond the public cloud, which isn’t always something platform teams are prepared for.

Edge + AI is turning fragmentation into a business risk

Today, edge and disconnected environments, such as air-gapped computing, have moved from niche use cases to mainstream constraints. When connectivity is intermittent or when latency matters, platform assumptions change. In these environments, reliability isn’t an IT metric. It’s a business outcome. Even minutes of downtime can cause major financial loss. It’s also a sign that data gravity is driving more pragmatic architectural conversations about the growing need to locate compute and data services closer to where data is generated.

AI raises the stakes further. If you’re collecting data at the edge, shipping it away for processing and pulling results back can be too slow and too expensive.

Our platform demands before betting on it

Before we’d bet on any platform, we’d ask a basic question: Can a single team operate both VM and Kubernetes environments without duplicating the entire organization? We’d insist on consistent governance: security controls and role-based access control (RBAC) should not fracture just because workloads are deployed differently.

We’d also look for cloud-like data services — object, file, block, and database capabilities — delivered quickly enough to keep developers moving toward their delivery targets, and designed to scale easily as application usage expands.

Then we would evaluate whether the platform helps reduce on-prem duplication. If it forces parallel networking, storage, and operational runbooks, the cost structure won’t improve.

Finally, we’d scrutinize lifecycle operations, including patching, upgrades, and maintenance, because “heroic” weekend work isn’t a sustainable strategy.

Dual native architecture is the pragmatic model

We use “dual native” to reject the binary choice. Enterprises need platforms that are both VM-native and container-native. Some workloads benefit from the operational efficiency of virtualization. Others are sensitive to latency or specialized hardware and are better served on bare metal. A one-size-fits-all mandate creates friction on both sides.

Dual native platform architecture isn’t just integration. It’s the one operational model that treats VMs and containers as first-class citizens. Teams no longer have to pick one architecture or stitch together separate stacks. In this model, organizations can keep mission-critical VM workloads running while building and scaling new cloud-native applications. Teams can maintain consistent management, governance, lifecycle operations, and cloud-like data services across VMs and bare metal servers across globally distributed infrastructure. 

NKP and NKP Metal as a dual native architecture

Nutanix Kubernetes Platform (NKP) solution with NKP Metal, which extends the Nutanix operating model and the NKP solution, supports Kubernetes deployments directly on bare-metal infrastructure. This solution provides unified Kubernetes operations, shared data services, centralized visibility, and automated bare-metal lifecycle management to support a dual native platform architecture.

Our approach with NKP starts with the premise that VM and bare-metal Kubernetes should operate under a consistent model rather than be split into separate toolchains and teams. To that end, a major focus has been on unified data services across deployment targets so the storage layer doesn’t become the breaking point when workloads span VMs and bare metal. We also purposefully centralize day-to-day operations and visibility across VMs and Containers in NKP so teams aren’t forced to manage two worlds with two separate management planes.

NKP Metal addresses lifecycle management, one of the biggest challenges of running bare metal at scale, including host OS setup, patching, and upgrades without resorting to late-night or holiday/weekend manual maintenance windows.

What’s next

Some things we know with confidence. VM workloads aren’t disappearing — the coexistence of VMs and containers will remain the operating model for many enterprises well into the next decade. Edge and AI workloads will likely continue to pull compute toward where data is generated, and budget pressure on duplicated infrastructure will likely only intensify.

What we don’t know is the pace. How quickly enterprises consolidate two platform teams into one depends on skills, internal politics, and licensing decisions, which vary widely from one enterprise to the next. Nobody can credibly predict a timeline there.

What we think is coming: AI inference at the edge will make bare metal a first-class deployment target rather than a special case, and platform teams will be judged less on which architecture they picked and more on whether developers can self-serve their own infrastructure, including data services, without opening a ticket.

The path forward is about building an operational foundation that accepts reality where VMs, containers, and bare metal coexist under a unified model.

The path forward is about building an operational foundation that accepts reality where VMs, containers, and bare metal coexist under a unified model. Enterprises that will thrive in this future are those adopting dual-native approaches that are ready for whatever comes next.  

The post “Just rewrite it”: What platform teams really think about modernization appeared first on The New Stack.

The path to artificial superintelligence

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy.

Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate patient care without a human making the decisions.

“The intelligence is already there. What is missing is the connective tissue that turns four strangers into one team,” explains Vijoy Pandey, senior vice president and general manager of Outshift by Cisco.

This “connective tissue” comes from adding a semantic layer—what Outshift calls the “Internet of Cognition”—that enables agents across domains to work together and, critically, “think” together through shared intent, context, and reasoning.

This semantic layer relies on a connectivity layer beneath it called the “Internet of Agents,” which allows autonomous agents to discover one another, prove identity, and exchange messages across domains.

When used together, they enable “the next step on the road to distributed artificial superintelligence,” says Pandey.

From solo silicon savants to the ‘Internet of Cognition’

For years, the AI industry has been focused on growth. Scaling vertically has led to bigger models, trained on more data with more compute. This has produced the reasoning capabilities that can be like a “brain” for AI agents, which can perceive, reason, and act in digital environments.

While vertical scaling can produce more capable agents perpetually, to enable agentic problem solving across different systems, companies, and platforms the next axis of scale must be horizontal, says Pandey.

Multi-agent systems are already being explored in areas like software engineering, drug discovery, and scientific simulations, but their performances so far have been underwhelming. One study finds a failure rate of between 41% and around 87% when evaluating seven open-source multi-agent systems.

“Connected agents handle coordinated action well; taking a task whose shape they have seen, divided and passed around,” Pandey explains. “What they cannot do is hold a goal in common and reason toward something none of them was trained to solve.”

“The gap is architectural, not a prompting problem,” Pandey adds. “Without the right coordination layer, naive multi-agent setups can perform worse than a single agent. The step change is that team of agents converging on its own, on a new problem, with no human stitching the seams.”

To reach this goal, Pandey says Outshift has built a connectivity layer called AGNTCY, an open-source project now under the Linux Foundation. AGNTCY allows agents across different systems, companies, and platforms to find each other, prove identity, and exchange messages through open, standardized protocols.

And, as Pandey explains, this allows the Internet of Cognition thesis to take a step further. It creates a semantic layer that allows agents to align goals (share intent), pool institutional knowledge and compound memory (share context), and make collective trade-offs (share reasoning).

Pandey likens this progression to that of humans: “For hundreds of thousands of years humans got individually smarter, and the gains died with each person who made them,” he explains. “Around 70,000 years ago that changed, when humans learned to share intent, build cumulative knowledge, and reason collectively. That is when scattered individuals became civilization.

“Agents are at the same threshold. We have built the silicon geniuses and given them agency. What they lack is the layer that let humans go collective,” he says.

First steps to distributed superintelligence

Enabling agents to work collectively rests on three pillars in the tech stack:

Shared intent through cognition state protocols: Cognition state protocols are the semantic handshake that allow agents to agree on a goal before they act and then negotiate toward it. Outshift has created an open-source coordination layer called Mycelium, which organizations can clone and use against their own agents.

“We found that unstructured groups reached a decision about a third of the time across 14 scenarios,” says Pandey, speaking about internal testing. “A coordination protocol that makes agents declare a goal, surface missing information, and resolve conflicts before acting raised that to 93%.”

Shared context through cognition fabric: A cognition fabric is a shared institutional memory and communication mesh that allows agent insight to compound over time rather than resetting each session. This policy-governed context layer solves the problem of “organizational amnesia,” says Pandey, by ensuring the baseline intelligence of the systems only ever goes up.

Shared reasoning through cognitive amplifiers and guardrail technologies: Two kinds of cognition engine can be used together to enable shared reasoning. Cognitive amplifiers speed up shared reasoning and modeling, and guardrail technologies (GATs) create security, cost, and compliance frameworks. Humans are active contributors to this layer, making judgment calls the system routes to them (rather than reviewing outputs after the fact).

Cognition sharing in multi-agent systems can create new risks, including unintended delegations, malicious prompt injections or memory poisoning, or over-privileged agents with access to permissions and data far beyond what their tasks require. Environment-specific controls are therefore needed to protect against unintended actions or consequences.

“Agents have human-like attributes but operate at machine speed and scale,” says Pandey. “Everything we built for twenty years—access control, identity, compliance—was built for humans or machines, not both.”

Continuous Agent Semantic Authorization (CASA)—an open-source reference implementation developed by Outshift—is a GAT that works to ensure agent actions remain securely aligned with the user’s original goal through a process of continuous authorization. It does this by reading what the agent is trying to accomplish then checking each tool request against that task.

In the case of a healthcare system, for example, an agent told to summarize a patient record may start by querying a whole database. This could lead to CASA denying the call, because the request no longer matches the task it was authorized for.

“Today’s controls are scoped to a role or a session not to the task so an agent granted a tool can use it for anything,” explains Pandey. “Roughly 90% of the time, an agent has no way to confirm it is even cleared for the job it was handed.”

Experimentation for cross-domain innovation

When horizontally scaling intelligence in the enterprise, businesses should begin by experimenting with one cross-functional workflow that spans three or four teams and currently needs a human authorizing the handoffs, Pandey advises.

“Stand it up as a small multi-agent system on open, interoperable infrastructure, with a measurable baseline,” he says. “Keep building bigger models, add the horizontal axis on top of them, and change what you measure. Track where one agent’s insight made another agent better—that is the signal the horizontal axis is working.”

By starting to experiment now with intent, context, and reasoning layers, organizations can get ahead of the curve. “The problems are open, and the infrastructure is still being written,” says Pandey. “This is the moment to build it.”

For more information on the Internet of Cognition, visit Outshift.com.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.




Closing the data loop in AI-driven drug discovery

Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage.

Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years—a phenomenon known as Eroom’s Law. Today, bringing a new drug to market takes an average of 10-15 years and costs anywhere from $1 billion to $2.5 billion, with failure rates upward of 90%.

AI has become the pharmaceutical industry’s biggest bet on bringing success rates up and timelines down. The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in development.

“The main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial,” says Paul Belcher, director of protein research strategy at global life sciences company Cytiva. “AI is one approach that drug companies hope will not only save time and compress timelines, but enable better quality candidates to reach the clinic.”

Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in lab systems.

AI brings efficiency to the lab

One of the most promising early-stage applications of AI in drug discovery is in hit identification. This involves screening libraries of molecular entities against a disease-related target, such as a protein, to find molecules that bind to it. A successful hit gives researchers a starting point for further testing and refinement, with the aim of eventually developing a viable drug.

Belcher has seen a shift from empirical screening to predictive design: Instead of physically screening libraries, drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing anything to research and development (R&D).

This means companies are no longer limited by how much they can physically screen to identify starting points. “AI does away with that,” says Belcher. “And it can help eliminate low-quality candidates before you have to physically test them, saving time and resources.”

What AI can’t do yet is reliably predict kinetics or developability of new compounds, says Belcher. This means every AI-generated candidate still needs to be validated in the lab.

Traditional screening workflows were built to identify hits at scale, not to profile large numbers of complex candidates in detail. This is placing more pressure on lab teams, who now have to test, characterize, and purify a growing volume of more diverse, AI-generated compounds.

“The current techniques used in hit identification can screen hundreds of thousands, sometimes millions of compounds, using binary or threshold-based techniques producing low-fidelity data—yes-or-no responses,” Belcher explains. “AI can increase the number of hits you get and potentially give you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits.”

Models need complete, quality data

As AI has accelerated demand for data-rich lab systems, it has also highlighted a fundamental need for better, more complete data.

Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. Because models have access to the same data, they all reach similar conclusions, with diminishing returns over time. Additionally, the datasets weren’t built with AI in mind, meaning they lack the structure, labeling, and diversity needed to keep models accurate and free of bias.

Publication bias reinforces the problem. “Most publicly available datasets and scientific publications focus exclusively on positive results,” says Belcher. “No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable.”

The data Belcher believes would markedly improve models—the failed experiments, the compounds that don’t bind—remains frustratingly difficult to come by. “We often joke that there should be a journal of negative data,” he says. “It’s often buried in lab notebooks, and it’s never used to inform or guide future research.”

This lack of negative data creates a fundamental problem: Without access to a broad range of data, models can’t be adequately trained to avoid bias. “In all machine learning applications, the model’s performance relies heavily on the quality and scope of the training data,” notes Belcher.

Fabrication has also become much easier with AI, compounding concerns around data integrity. Take Western blots, for example. These are part of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. This was back in 2016, before generative AI made fabrication trivial.

“Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences,” says Belcher. “There needs to be more tools to verify that data is not manipulated.”

Some vendors are starting to tackle this challenge. Belcher points to solutions like Cytiva’s Image Integrity Checker, for instance, which uses secure hash algorithms—the same technology used in blockchain—to detect whether scientific images have been tampered with. “We’re starting to see a lot of interest from publishing houses that want to adopt this as standard because it’s a quick way to ensure that what gets published in the literature is genuine,” he adds.

Autonomous labs could accelerate breakthroughs

Belcher describes the future state of drug discovery as fully autonomous labs that run with minimal human intervention. Foundational to this vision is consistency in data and infrastructure.

These AI-driven dark labs, or labs-in-the-loop, operate around the clock. They cycle through prediction, testing, and optimization, and then feed results back into AI models to guide the next round of experiments. This can improve the success rates of drug candidates entering clinical trials, says Belcher. Better starting points, combined with more rounds of optimization, should result in better candidates with fewer liabilities reaching the clinic.

But automating a lab depends heavily on integration. That means interoperable systems, highly structured and comprehensive datasets, and information flowing easily in and out. Most labs aren’t there yet. “Today, a lot of the instruments in labs are standalone,” Belcher notes. “You can have the best technology in the world, but if it’s a closed ecosystem—if the user can’t get the data out—it doesn’t do any good.”

An integrated infrastructure can enable labs to generate FAIR (findable, accessible, interoperable, and reusable) data at scale. This would not only inform individual lab reports, but could also train subsequent generations of AI models, effectively closing the loop between the computational, AI-driven dry lab and the physical wet lab.

“Our goal is to help scientists and researchers accelerate their breakthroughs and make that future state of autonomous labs a real possibility,” says Belcher. “We want to help them generate reliable data, simplify workflows in discovery, and hopefully enable what they’re working on to become tomorrow’s life-changing therapies, faster and with greater confidence.”

On costs and what comes next

AI-driven drug discovery is still in its early days. Notably, no drug discovered primarily through AI-driven design has yet received full FDA approval—although Belcher expects that to change in the next two to three years.

How big of an impact could AI eventually have on drug discovery? “The holy grail would be full in silico prediction of efficacy and toxicity, eliminating the need for the vast majority of physical wet lab work,” says Belcher. But there are many barriers to this beyond the maturity of the models, including regulatory hurdles and cost challenges.

A Stanford study found that the cost of training frontier AI models has more than doubled every year since 2016, adding more financial pressure to a sector already defined by exceptionally high R&D spend.

Belcher acknowledges the tension, but remains optimistic about what’s ahead. “I think we’ll get to a point where there’s a balance between AI and wet work, from a cost perspective and a risk perspective,” he says. “As long as the cost of compute doesn’t ever outweigh the cost of clinical development, I think AI is going to be an advantage.”

Learn more about how Cytiva is using faster discovery to reshape protein purification workflows.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Building the enterprise environment for agentic AI

For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the ability to predictably plan and scale agents.

To better understand some of these dependencies, Intel performed thousands of agentic AI workload experiments. Our initial findings create and support five practical lessons for enterprise leaders:

  1. Agentic AI is a larger systems problem, not just one of inference.
  2. The majority of existing agentic AI harnesses are limited and do not measure overall system performance.
  3. Plan capacity is done using agents per virtual CPU (vCPU) density, not agent count.
  4. Monitor agent task latency, not just average CPU utilization.
  5. Default to scale-out for systems hosting agents. Reserve scale-up for workloads with heavier per-agent compute or architectural constraints.

Beyond inference: Agents as workflow automation

Agentic AI is more than LLM inference. Its enterprise value depends on the full system, task orchestration, data access, tool execution, latency management, governance, and scalable infrastructure. An agent is a goal-driven automated enterprise workflow process: It plans a multi-step task, calls tools, reads results, and retries when something fails. Enterprise agents are therefore not just an inference problem; they are a systems problem.

Defining what good looks like

Most agentic AI metrics focus on evaluating the LLM used. Platform teams also need to know how long the tasks take, how many agents a fleet can support, what users experience at the end of the execution process, and how costs change as more agents work simultaneously.

A more useful enterprise view looks at six metrics:

  1. Task success rate
  2. Cost per task
  3. Time per task
  4. Task throughput
  5. Agent density (agents per vCPU)
  6. Latency

Together, these answer the questions enterprise AI operators care about: Is the system performing as expected? How many agents can the system sustain? How should it scale to support more agents?

Building on solid foundations

To gain a deeper insight into agentic AI workload performance, Intel extended Terminal-Bench, an open source benchmarking harness for evaluating AI agents with profiling, telemetry, and replay capabilities. This made it possible to understand where the agents spent time beyond LLM inference.

The benchmark extension used a deterministic record-replay of LLM responses to separate agent performance from LLM variability. LLM responses were recorded once and replayed identically across runs, reducing run-to-run variance and creating a more reliable basis for comparison.

The Terminal-Bench task mix used was intentionally broad. It included compilation, testing, database operations, Boolean logic, interpretation, ray tracing, compression, linear algebra, video transcoding, and machine learning training. That wide variety made the findings more relevant to real enterprise environments.

Agentic AI in three dimensions

Deploying agentic AI should be approached in three phases:

Plan in terms of agent density, not agent count: The first sizing rule is to normalize agent count by available compute. Agent density, measured as agents per vCPU, is the leading signal for saturation. For example, 10 agents on an 8-vCPU system and 20 agents on a 16-vCPU system behave similarly if the density is the same. This gives architects a portable way to compare capacity across instance sizes and processor generations.

The right density also depends on the business goal. Interactive copilots and user-facing assistants should favor lower density because response time matters. Batch workloads such as IT workflows can often run at higher density. This gives teams a practical way to tune fleets around service-level objectives and total cost of ownership.

Agentic AI requires a new form of observability: Average compute (CPU) utilization is a weak primary performance monitoring signal for agentic workloads. Agents often alternate between waiting for model responses and then doing short bursts of compute-intensive work. Because of that “bursty” pattern, average utilization can look acceptable even when those bursts are creating queues and slowing down the user experience. Task latency (P95) is a better leading metric. It shows when workflows are starting to wait, even before average task duration meaningfully degrades. A practical operating model is to alert on P95 latency first, then confirm the issue by looking at sustained task duration.

Scale out by default: Scaling out adds more systems, increasing total agent capacity, while scaling up adds cores or memory to a single system for agents with heavier compute bursts.

Our testing data showed that scale-out is usually the better default. That aligns with the fact that agents are typically semi-independent and have modest per-agent bursts, it improves overall performance, supports high availability, often lowers cost, and makes it easier to preserve the target agents-per-vCPU ratio as the platform grows.

Scale up when agents require heavier parallel compute, shared state limits partitioning, memory locality matters, or licensing constraints apply.

Consider business implications: Where will agentic AI create business value first? The organizations getting production-grade results are wrapping an automation layer around workflows that already have codified rules and measurable service levels: code creation, regression test farms, ticket triaging, market analysis, and security review.

The ideal enterprise persona for agentic AI is therefore not the experimental user chasing novelty; it is the accountable leader who must improve cycle time and productivity, protect service quality, enforce policy, and scale adoption with cost in mind. 

Agentic AI’s value comes from helping businesses complete real work across teams, systems, data, and processes. For enterprises, the priority is not just better model performance; it is creating a reliable environment where AI agents can support business workflows, improve productivity, operate within governance requirements, and scale as adoption grows.

In practice, success with agentic AI depends on the right foundation to deliver consistent outcomes, manage cost, maintain control, and move confidently from pilots to production.

This content was produced by Intel. It was not written by MIT Technology Review’s editorial staff.



How AI helps scientists design the next generation of medicines

Designing and developing a new medicine is an expensive, failure-prone scientific challenge. A new drug can take many years to develop, at the cost of a significant investment. And even then, most possible candidates never reach the patient. For biologic medicines, therapies made from engineered proteins rather than synthetic chemistry (which are often used to treat conditions across most major acute and chronic diseases), the complexity is even greater.

Scientists explore vast quantities of possible molecules, looking for the rare few that will bind to the right target, remain stable in the human body, and be manufacturable at scale. Today, AI is speeding up these processes and has quickly become a core part of the infrastructure in pharmaceutical R&D.

AI-assisted design is a growing part of how biologic drug candidates are developed, and companies like AstraZeneca are actively building its engineering teams to push this further. “Everything we do, whether it’s design, make, test, or analyze, is now computationally enhanced,” says Puja Sapra, senior vice president and head of R&D biologics engineering and oncology targeted discovery at AstraZeneca. “The cycle times are getting shorter while productivity and innovation increase.”

Sapra explains that AstraZeneca’s approach follows a build-measure-learn loop. AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed. Scientists then focus lab resources only on the top-ranked candidates. This leads to a tighter feedback cycle with fewer dead ends, faster iteration, and the ability to go after disease targets that were previously considered untreatable by medicine. Because the number of possible molecular combinations far exceeds what any human team can systematically explore, using AI to narrow and refine the options for testing has become a major focus in biologics drug design.

Navigating complex drug design problems

Beyond accelerating timelines, AI is also being applied to the discovery of entirely new classes of medicines. Traditional biologics typically target one disease pathway. The next generation of drugs can hit multiple targets simultaneously or precisely deliver therapeutic payloads to specific cells. Achieving this requires optimization across many variables at once. Looking ahead AI-driven models could help design these increasingly complex, multi-specific biologics, explains Puja Sapra. “For example,” she continues, “such models could help identify which two or three targets to prioritize based on the underlying biology, then optimize across multiple parameters to balance a molecule’s potency, stability, manufacturability, and safety.” “Drugging the undruggable is becoming a reality,” Sapra says. “These technologies will eventually enable us to develop medicines against targets once thought impossible to reach. The potential for benefit to patients is remarkable.”

The data moat

McKinsey estimates that generative AI, combined with other computational tools, could cut drug discovery timelines by as much as 50%. But every AI model is only as good as its training data. In drug discovery, that means ample quantities of high-quality biological data. Experiments can provide a rich source of such data. Whether they succeed or fail, each experiment generates a signal about what does and does not work.

“Data is our differentiator,” says Sapra, explaining how the company’s datasets are proprietary and multimodal and include molecular structures, binding measurements, safety profiles, and manufacturing outcomes. “We’ve built an intentionally diverse portfolio across multiple disease areas and drug types. All of that data empowers us to fine-tune frontier AI models with richer, more representative training sets.” She continues, “Further, we have invested in deep screening technologies to generate additional datasets required in volume to constantly refine and validate our models.”

Building an autonomous discovery engine

To bring all of that data together in one place, AstraZeneca is building what it calls a “lab of the future” facility in Kendall Square, Cambridge, Massachusetts where AI and robotic automation will be able to form a continuous, closed-loop discovery system. “Where a self-driving car uses sensors and models to navigate its environment, this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data,” explains Sapra. That data feeds directly back into the models, accelerating each subsequent cycle.

“Throughout, scientists will remain central to the process, providing the oversight, judgement, and strategic direction that ensure outputs are explainable, tolerable, and directed toward potential patient benefit,” she adds.

Eventually, automated high-throughput systems will be able to make and evaluate thousands of molecular interactions on a weekly basis. “This will generate AI-ready data at a scale that traditional workflows cannot match,” Sapra says. “Robotic sample handling, automated quality checks, and integrated data pipelines also have the potential to help accelerate early drug development timelines significantly.”

The next frontier: Generating medicines from scratch

Ultimately, Sapra says, the end-state vision for AI in biologic drug discovery is what the field calls “de novo” design. For this, the goal is for AI to generate entirely new protein sequences that precisely fit the desired drug properties. This includes designing the structure, predicting safety, how it will behave in the body and how to make it manufacturable.

“The field is making great progress toward a completely AI-generated biologic, designed from scratch all the way to a clinical candidate,” Sapra says. “As we continue to leverage frontier models and fine-tune them with the right datasets, we bring ourselves closer to this reality. I believe it will come. It’s a matter of time.”

Several key elements are needed to reach this point, however. First is richer and more standardized training data across the industry. Second, robust evaluation benchmarks for AI-generated candidates. And third, teams that know how to work at the intersection of machine learning and biology. Of all the prerequisites, however, safety prediction may be the most consequential, and perhaps the least discussed, Sapra says.

“One of the hardest problems in de novo design is predicting whether a computationally generated molecule will be safe in the human body,” Sapra explains. AstraZeneca is tackling this with what amounts to virtual clinical trials. These are advanced cell systems and micro-scale organ models that function as physical testbeds, paired with AI that learns from their outputs.

 “These systems have the potential to generate enhanced biological signals without traditional testing bottlenecks, and they’re a critical missing piece in closing the loop between AI-generated designs and clinical-ready candidates,” Sapra adds.

A shift currently underway is the move toward agentic AI systems that can simultaneously generate molecule candidates and predict how efficacious and safe they are likely to be. These autonomous workflows can connect disease-level insights directly to molecule design, bridging what were previously separate data silos. “The complexity of the biology goes hand-in-hand with the design of the molecule,” summarizes Sapra.

Human talent unlocks AI potential

The transformation underway in biologics is not just about technology. “With more autonomous systems, human oversight remains at the heart of this approach—ensuring explainable and ethical AI for the benefit of patients,” says Sapra.

For scientists, working with AI is a collaborative process. “Scientists will work hand-in-hand with these model systems,” she says. “There will be a world where models will design molecules, then scientists will work with the systems to test those molecules and put all that data together.” Through this process of human checks, balances, and judgement calls, the models will evolve and constantly improve, ultimately with potential to benefit patients.

For engineers, designing and building effective systems ready for human-AI collaboration will mean ensuring high levels of model transparency and explainability. According to Sapra, AstraZeneca’s engineering teams include data scientists, automation specialists, and AI engineers, who are developing systems that act as “thinking partners” rather than black boxes. “Engineers are designing systems that generate, validate, and learn at speed. And the problems are genuinely hard: Multimodal data fusion, closed-loop optimization, uncertainty quantification, and interpretability at the point of clinical decision-making,” she adds.

In taking on such technically demanding challenges, engineers and scientists have the opportunity to contribute to the research and development of potentially life-changing treatments for many diseases, says Sapra. “The biologic medicines we can develop today, and those we’ll design tomorrow, depend on combining world-class AI and engineering talent with deep scientific expertise.”

This article has been initiated and funded by AstraZeneca.  Z4-85058, July 2026.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Advancing next-gen AI with materials science innovation

The conversation about AI often centers on algorithms, computing power, or huge investments in new semiconductor fabrication plants and hyperscale data centers. But beneath each of these advances is another layer of innovation that makes them possible: advanced materials.

Every new generation of AI technology demands more processing power, more memory, greater energy efficiency, and higher reliability. Every increase in computing performance increases the physical demands placed on the systems that make and run AI.

Delivering these gains depends not only on advances in chip design and system architecture, but on advances in the materials that enable them to perform under extreme conditions.

As AI continues to push the physical limits of semiconductors and data center infrastructure, advanced materials are no longer simply supporting innovation in this area; they are defining the limits of what is possible.

Performance first

Advanced materials exist to solve performance challenges. As AI raises the bar, these challenges are becoming more demanding.

Manufacturing a semiconductor chip today requires thousands of tightly controlled process steps, with almost no room for error. Tiny variations in temperature or chemical instability can create defects that reduce yield and drive up manufacturing costs. With every new generation of semiconductor chips, manufacturers seek advanced materials that can deliver greater purity, higher chemical and plasma resistance, and better stability under increasingly harsh operating conditions.

These are familiar engineering challenges being pushed to new extremes. And it’s here that materials innovation makes the difference with continuous advances in polymers, elastomers, specialty fluids, and other advanced materials that make each new generation of technology possible.

For materials companies, it’s not about reinventing semiconductor manufacturing but about ensuring the materials supporting the industry continue to evolve alongside it. This same principle applies beyond the semiconductor fabrication floor. As AI workloads become more demanding, the physical infrastructure that powers them is evolving rapidly.

Increasing computing density is transforming data center design, driving the need for more sophisticated thermal management, higher-voltage power architectures, increased data storage, and faster, more reliable data transmission. Every part of the system is under greater pressure, from cooling and power management to critical electronic components, such as connectors, capacitors, and hard disk drives.

At Syensqo, we’re building on our expertise in electronic and electrical components, along with insights from other markets, to meet these emerging needs.

For example, as data centers shift to higher-voltage architectures and greater power density, many of the materials challenges we face closely mirror those of electric vehicles. Fluid-circulation know-how from semiconductor and automotive coolant systems, for instance, can be adapted to direct liquid-cooling designs for AI servers. By transferring knowledge across markets, we can accelerate new power and thermal management solutions while supporting the reliability required by next-generation AI infrastructure.

Whether we’re talking about semiconductor fabrication or hyperscale server farms, the challenge for materials science companies is the same: enabling greater performance without compromising reliability.

A new definition of what performance means

While performance remains the first priority, the way performance is defined is changing.

In addition to meeting the increasingly demanding technical requirements of next-generation semiconductors and data centers, there is now an expectation that these materials are developed and manufactured more responsibly.

Perfluoroelastomers, for example, are used to seal semiconductor manufacturing equipment. These materials operate under extreme temperatures, aggressive plasma, and highly reactive chemicals.

To make the process more sustainable, at Syensqo, our next generation of perfluoroelastomers use a fluorosurfactant-free manufacturing process. Our goal was to make a better-performing material, produced in a better way, ensuring manufacturers no longer have to choose between higher performance and a more responsible way of producing the materials that enable it.

This approach reflects a broader reality across the industry.

New materials aren’t adopted simply because they are new. Qualification can take years, and manufacturers only make changes when a material solves a genuine engineering challenge or enables new technology.

Performance remains the price of entry. The difference today is that the definition of performance has expanded. Success increasingly depends on delivering technical excellence through more responsible manufacturing from the outset.

Accelerating the pace of discovery

As the performance bar rises, the way we innovate must evolve with it.

Developing advanced materials has traditionally involved a lengthy process of hypothesis, synthesis, testing, and iteration. While this process remains unchanged, new digital tools are helping researchers move through these cycles faster. By helping researchers identify the most promising candidates earlier, AI can reduce the number of physical experiments required and accelerate the earliest stages of materials discovery.

AI isn’t replacing scientific expertise. It’s helping scientists apply that expertise more effectively, allowing them to spend less time searching for answers and more time solving the industry’s toughest challenges.

At Syensqo, we’re putting this approach into practice through use of several AI tools, including the Microsoft Discovery platform, which are helping researchers identify and evaluate promising molecular candidates for next-generation heat transfer fluids, used in semiconductor manufacturing and data centers.

AI helps our researchers rapidly identify and evaluate promising molecular candidates based on the properties they need to achieve. This allows us to focus laboratory work where it has the greatest potential to deliver results, accelerating discovery and reducing the time needed to turn promising materials into solutions customers can qualify and deploy.

The journey from laboratory discovery to a qualified material will always require scientific expertise, rigorous testing, and close collaboration with customers. But by accelerating the earliest stages of discovery, AI can help materials innovation keep pace with the evolving needs of industries such as semiconductors, electronics, and data centers.

Progress is earned

The future of artificial intelligence will depend on better algorithms, more powerful chips, and larger computing infrastructure. But sustaining that progress will also require advances in the materials that make those technologies possible.

Whether in semiconductor manufacturing or AI infrastructure, progress is earned. Every new generation of technologies raises the bar, and every new material must prove it can deliver the performance, reliability, and efficiency needed before it earns its place.

For materials companies, that remains both the challenge and the opportunity.

This content was produced by Syensqo. It was not written by MIT Technology Review’s editorial staff.

Arm and Google offer a smarter option to run agentic AI workloads

Warp speed light streaks radiating outward on blue background

As enterprise leaders start deploying agentic workflows, they must establish the infrastructure to build and run them, one capable of fluidly routing a diverse set of workloads across the most efficient compute resources.

This requires the ability to manage heterogeneous infrastructure, utilizing high-performance accelerators for large-scale training and inference, and utilizing CPUs for the critical orchestration layer of agentic AI. As autonomous agents become more prevalent, CPUs are ideally suited for managing agent state, semantic routing, tool selection, and spinning up secure, isolated sandboxes to safely execute untrusted generated code.

The Google Axion advantage

Google Cloud, with its workload-optimized Compute Engine portfolio, which includes general-purpose and specialized offerings, shines in addressing this need.

Google Axion processors within this portfolio comprise a family of custom Arm processors engineered for performance, efficiency, and versatility, with a feature set that supports general-purpose workloads, CPU-based AI workloads, and other specialized tasks requiring Arm-native compatibility and direct hardware access.

Axion is Google’s first custom Arm-based server CPU, introduced in April 2024. It is designed specifically for hyperscale cloud and AI-era data center workloads. 

Axion also leverages more than a decade of Google’s custom silicon innovation. This enables Google to more readily incorporate customer feedback into chip designs and address the more general, though complex, needs of CPUs. 

Matching workload type to the processor

Bhumik Patel, Director of Software Ecosystem Development at Arm, says the key to all of this is to match the workload type as closely as possible to computing capacity. CPU-powered cloud instances are a practical option for certain AI workloads, particularly those with smaller datasets or less complex models. 

“Agentic tasks such as orchestrating, talking to APIs, and memory management are all ones CPUs are good at, so it’s a distributed and concurrent AI workload,” Patel tells The New Stack. Intelligent workload-processing apportionment makes agentic AI more cost-effective and efficient than running all workloads on a single compute type.

This efficiency is quantifiable. The Google Kubernetes Engine Agent Sandbox running on Google Axion N4A provides up to 30% better price performance than the next hyperscale cloud provider, says Google’s Mo Farhat, Axion Group Product Manager. The GKE Sandbox is an open-source Kubernetes-native primitive designed to execute untrusted AI-generated code safely. 

“Agentic tasks such as orchestrating, talking to APIs, and memory management are all ones CPUs are good at, so it’s a distributed and concurrent AI workload.”

Intelligent workload decoupling makes agentic AI significantly more cost-effective. Google Cloud’s fluid computing foundation enables engineering teams to reserve specialized accelerators strictly for heavy reasoning and generative workloads, while leveraging Axion CPUs for high-concurrency orchestration and context management.

Secure execution with the GKE Agent Sandbox

As agents begin to generate and execute dynamic code autonomously, security is non-negotiable. Running AI-generated code directly in a standard cluster poses severe security risks, as untrusted code could potentially access other apps or the underlying cluster node.

The Google Kubernetes Engine (GKE) Agent Sandbox resolves this by providing an isolated environment for safely executing untrusted code. Running on Axion-powered N4A instances, the sandbox provides up to 30% better price performance than comparable workloads on other hyperscalers.

The vertical stack isolates sensitive tasks at the kernel level with sub-second latency.

The vertical stack isolates sensitive tasks at the kernel level with sub-second latency.  GKE Agent Sandbox natively supports gVisor (an open-source application kernel developed by Google that acts as a secure sandbox for containers) and default-deny Kubernetes network policy. Agent Sandbox provides pluggable interfaces for open-source sandboxes, such as Kata Containers, enabling users to customize their kernel isolation. 

Powered by gVisor technologies with software support from Arm’s architecture, the sandboxes intercept and validate system calls before they reach the host kernel. These isolated execution environments enable deployment of autonomous systems at scale without sacrificing performance or operational agility.

To manage resources efficiently when agents sit idle, GKE Pod snapshots allow users to save and restore the exact process state of sandboxed environments. This functionality provides four major architectural benefits:

  • Fast startup: Reduces sandbox startup time by restoring from a pre-warmed snapshot rather than initializing from scratch.
  • Long-running agents: Pauses sandboxes that take a long time to run and resumes them later—or moves them across nodes—without losing progress.
  • Stateful workloads: Persist an agent’s context, such as conversation history or intermediate calculations.
  • Reproducibility: Captures a specific state to use as a baseline for spinning up multiple new sandboxes.

Getting started

As token generation, autonomous workflows, and continuous agent interactions grow exponentially, relying exclusively on accelerator-backed stacks for every task will become financially and architecturally unsustainable.

The combination of CPU and accelerator execution accounts for bursts in agent activity and unpredictable demand spikes by eliminating the inference tax. Google Cloud’s full-stack advantage enables organizations to deploy the right machine for the job. 

By using Google Axion and GKE Agent Sandbox, builders can optimize total cost of ownership and security while maintaining the performance required for AI agents.

Learn more about Google Axion.

The post Arm and Google offer a smarter option to run agentic AI workloads appeared first on The New Stack.

Building a Foundation Stack for General-Purpose Robots

13 July 2026 at 10:19


This article is brought to you by X Square Robot.

Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the field does not yet agree on what it is.

X Square Robot, a Chinese embodied-AI company, has made an unusually explicit bet. It argues that the recipe is an integrated stack, spanning the data a robot learns from, a world model for predicting changes in the physical world, and an action model that brings together perception, planning, reasoning, and decision-making to generate executable robot behavior. The company also believes that the stack should be built and released in the open.

X Square Robot shares its vision of bringing robots into real homes.X Square Robot

X Square Robot’s embodied AI stack

What holds the stack together is a small set of principles rather than a single overarching model.

  • The first is that the basic unit of robot data is an interaction, not a trajectory; a demonstration is successful only if it changes the world as intended, not simply because the joints moved.
  • The second is that pretraining should yield usable capability, not just an initialization for later fine-tuning.
  • The third is that behavior should be modeled around physical events rather than fixed slices of time.

These principles make the layers interdependent, since the same robot-free data that trains the action model is also structured to feed the world model. It is worth being precise, though. The company describes the world model and the action model as complementary but independent model families that share a code base. Both sit within its broader World Unified Model, which it has presented as an architecture for training vision, language, action, and physical prediction together.

Robot learning data: Engineering for quality and cost, not scale

For the X Square Robot team, one of the biggest constraints on general-purpose robots is the cost and quality of interaction data, not the number of parameters. To address that, the company built its Universal Manipulation Interface (UMI) data collection system, QUANXTA Zero Series. It works by collecting demonstrations from people wearing a rig with dual grippers rather than teleoperating a robot. This approach is not itself new, and builds on established methods for robot-free data capture. What sets it apart are two engineering choices.

Person using VR headset and handheld controllers to teleoperate a dishwashing robot system X Square Robot emphasizes data quality control, recording trajectories and replaying them on a real robot, with only those that actually complete the task counted as valid.X Square Robot

The first is quality control, and it is the most distinctive part. Rather than accepting recorded trajectories as they are, the system runs a closed inspection loop, and its notable step is physical playback. A sample of trajectories is replayed on the real robot, and only those that actually complete the task count as valid. That makes the validity rate a measured quantity rather than an assumption. For example, a gripper that closes a fraction of a second too early still looks like a grasp in the data, yet it has pushed the object away, so it shouldn’t be classified as valid. A smaller clean dataset can be worth more than a larger noisy one.

The second choice is how lower-cost human data and scarce robot data are combined. The company pretrains on a large volume of robot-free demonstrations to build general representations, then adds a small amount of real-robot data as an anchor to the specific machine’s dynamics. It reports that this reaches performance comparable to an all-robot dataset at roughly a 20-fold lower cost of collection, driven mainly by how much cheaper the wearable rig is than a teleoperation setup.

The resulting dataset is deliberately model-agnostic, formatted to feed both action models and world models. The caveat is that the strongest results are measured on the company’s own robots and data-collection pipelines. Broader independent testing will help confirm and extend these promising results across a wider range of settings.

A world model organized around events

In developing its world model, called WALL-WM, X Square Robot took a differentiated approach. Most action models predict a fixed-length chunk of motion from the current image and instruction. That is convenient, but it segments behavior into fixed-duration windows, so the boundaries fall where elapsed time dictates rather than where one action ends and the next begins. WALL-WM instead treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.

Collage of robot arms manipulating kitchen objects with charts of multimodal AI performance X Square Robot’s world model, called WALL-WM, treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.X Square Robot

WALL-WM’s design reflects a specific concern about not discarding what large video models already know. To achieve that, a text-to-video model is coupled to a freshly initialized action network that reads from the video features without overwriting them, which preserves the visual prior. From that one process, it offers two modes. An event mode runs in variable-length segments and suits reasoning over long horizons, while a fixed-length mode produces the steady, real-time output a controller needs. That places WALL-WM between mainstream chunk-based action models and pure video world models, keeping the predictive character of a world model while still yielding executable control.

In a series of experiments, the company relied on a generalization test that is more specific than most. A model trained on a limited dataset was evaluated on long-horizon tasks in unseen settings and, on the company’s real-robot benchmark, reportedly outscored baselines that had been fine-tuned on related data. That is a meaningful result if it holds. For now, it is measured on the company’s own benchmark. With the code now being released, the broader community will have the opportunity to test, reproduce, and build on them across more settings.

A policy that runs before fine-tuning, and action tokens with meaning

The action layer carries two connected ideas. The first is a requirement the company sets for itself with Wall-OSS-0.5, its vision-language-action model: The pretrained model should run on a real robot before any task-specific fine-tuning.

The interest is less in the scores than in the design behind them. The model trains three objectives together, namely discrete action tokens, language grounding, and continuous action generation. And it keeps gradients flowing through all of them rather than freezing parts of the network as some rival designs do. It’s also a more strict method, since it reports untuned behavior such as approaching, grasping, and recovering, including on a deformable task held out of training.

Dashboard of robot training metrics with charts and photos of a robot sorting objects As part of X Square Robot’s Wall-OSS-0.5 vision-language-action model design, the pretrained model should run on a real robot before any task-specific fine-tuning. X Square Robot

The second idea is the action interface itself, called X-Tokenizer. Most systems that turn continuous motion into discrete tokens produce codes that the language model cannot interpret. X-Tokenizer reframes tokenization as learning a semantic interface, so that the top-level code stands for the intent of a motion while lower-level codes carry finer detail, all aligned with the language model’s own features.

A useful consequence is stability. Adding noise to an action barely moves the intent code, which is what lets one tokenizer to be reused across robots without re-tuning. The tokenizer inside the production action model is a related variant of this approach. Together, the two ideas give the action layer something rather powerful: capability that transfers.

The future of embodied AI stacks

X Square Robot is betting that its unique approach combining three layers, each specialized in solving a key part of the problem, will stand out from other embodied AI stacks. The physical-playback step that grounds data quality is uncommon and sensible. The reframing of world modeling around events, with one backbone serving both reasoning and control, is a genuinely distinct approach. And the pairing of a deployable pretraining standard with a tokenizer designed as a semantic interface gives the action layer unusual coherence.

X Square Robot’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.

The next phase will bring broader validation. Much of the current evidence comes from X Square’s own robots and benchmarks. With the world model code now being made public, and as the community begins to test, reproduce, and build on the work, the reported capabilities will be tested across more robots, tasks, and settings.

X Square Robot’s recent funding rounds reflect similar confidence. The company’s valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.

What’s next for X Square Robot

To learn more about its future plans, the following Q&A with the X Square Robot team further explores the company’s technology, strategy, and vision.

What made now the right moment, technically, to commit to this stack? What recently became possible that wasn’t possible a couple of years ago?

It is not one breakthrough but several trends maturing together. Foundation models gave us a shared representation across vision, language, and action, so we can model what a robot sees, what it is asked to do, and how its actions change the world in one framework, rather than as separate perception, planning, and control modules.

Compute and infrastructure are finally sufficient for large-scale pretraining over long-horizon, multi-embodiment data. Just as importantly, we realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical. The useful question is no longer how to predict a few seconds of video, but how to understand the ways actions change objects, contacts, and task states. Two years ago these ingredients existed separately. Today they are mature enough to work as one system.

“We realized that data, not model size, is the real bottleneck for general robots—what is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical.”

Your data system captures demonstrations with a wearable VR rig and custom grippers rather than teleoperating robots. What was wrong with standard teleoperation?

Teleoperation is built around controlling the robot. It forces the operator to work within the machine’s kinematics, latency, and viewpoint, and the resulting demonstrations are slower, stiffer, and less diverse. We built our system around capturing human skill instead. Manipulation is really about contact, timing, finger coordination, and recovery, not just the path the hand takes, and a wearable rig records those before the behavior is compressed onto one particular robot. It also breaks teleoperation’s expensive scaling law, in which every demonstration needs a robot.

People can generate rich data independently of any robot, and the crucial property is that those demonstrations can still be replayed and executed on a physical robot through the model. Mobility is convenient, but that replay is the real point, because it is what lets the same data be reused across different platforms.

Robot and person loading a washing machine together in a modern laundry room. In X Square Robot’s approach, demonstrations can be replayed and executed on a physical robot through the AI model, allowing the same data to be reused across different platforms.X Square Robot

X Square Robot reports that its pipeline has roughly an 85 percent data-validity rate. Why is quality control such an underrated bottleneck?

Because errors in robot data are far more expensive than in language data. A small timing or contact error can change what a demonstration means. If a gripper closes a fraction of a second too early, the motion still looks like a grasp, but physically it has pushed the object away. A dataset that mixes failures and accidental successes teaches ambiguity, not skill, because the real unit is the interaction, not the trajectory.

So we run automated inspection, kinematic checks, and physical replay, where we play a sample of trajectories back on the real robot and count only the ones that actually complete the task. Data quality sets the ceiling on how good a policy can be. In our experience a smaller, cleaner dataset often beats a much larger, noisier one, which is why we treat quality control as part of the model, not a preprocessing afterthought.

The model runs in both “event mode” and “chunk mode.” When does each matter?

Both matter, for different reasons. The physical world changes through events—when contact occurs, a grasp forms, or an object slips—not in fixed-frame windows. Event mode concentrates the model’s attention on those moments, and it matters most for long-horizon tasks, like clearing a table, where progress is a sequence of semantic events rather than a smooth stream. It runs in variable-length segments that follow the task rather than a clock. Chunk mode matters for deployment. Real controllers need a stable, real-time interface, and fixed-length chunks integrate cleanly with existing control systems.

We organize learning around events in the first place because a fixed window can split one motion in half or merge two together, which turns training into short-horizon pattern matching and weakens the model on long tasks. So the world model’s job is to connect event-level understanding, which is where the reasoning happens, with a fixed-length output a real robot can actually run.

Why make “deployable before fine-tuning” the criterion?

Pretraining should produce capability, not just a good starting point. If a model is only useful after heavy fine-tuning, then most of the intelligence still lives in the downstream supervision, not in the foundation model. Deployable before fine-tuning is a more honest test of what pretraining actually learned. A well-pretrained robot should already know how to approach, grasp, move, avoid obstacles, and correct itself. Fine-tuning should adapt it to a specific task or robot, not create the ability from nothing. It is also a practical requirement. A robot in a home or a workplace shouldn’t need a brand-new dataset and a new policy every time the task changes, so a foundation model that already carries general skill, and some ability to recover, is the minimum bar for something genuinely useful in the real world.

What is the most challenging part of cross-embodiment learning?

Robots differ in control frequency, delay, compliance, sensing precision, and contact dynamics, so the same instruction can require different action decompositions and recovery strategies, and a behavior that works on one arm cannot simply be copied to another. Cross-embodiment learning needs an intermediate abstraction, lower than language but higher than joint angles: how you approach an object, how you make contact, how you apply force, and how you recover from a mistake.

When we say cross-embodiment, the main capability we mean is multi-embodiment generalization: transferring across robots, training on many embodiments at once, and adapting to different kinematics. Human-to-robot transfer and other techniques are specific approaches to that goal.

“A robot in a home or workplace shouldn’t need a new dataset and policy every time the task changes. A useful foundation model should already carry general skills and the ability to recover.”

What would you most like to see other researchers attempt to reproduce or stress-test?

Three things, above all. Whether event-level representations really generalize beyond our own datasets, across more tasks, scenes, objects, embodiments, and failure conditions. Whether pretraining stays effective on robots the model never saw during training, or whether its capability is still too tightly coupled to what it has already seen. And whether real-robot evaluation can become a shared language for the field, so that we compare not just success rates but the reasons systems fail, where an instruction was misread, where perception broke down, or where recovery fell short. Robotics has been driven too often by impressive demonstrations, and real progress comes from results that are reproducible and diagnosable.

What capability is still missing before robots become dependable in homes?

Benchmarks measure competence, like whether a model can finish a task. Homes demand reliability, safe and consistent operation over time in a place that changes every day, with objects moving, instructions that are vague, and people interrupting. The missing piece is not a higher one-time success rate: it is robust recovery. A dependable home robot has to know when it is uncertain, when to slow down, when to ask for help, and how to bring the world back to a safe state after it drops something or misunderstands a request.

In a real home, failure recovery matters more than raw success, because the home does not reset itself. Homes also demand careful personalization, learning a household’s routines and preferences over time, with safety and trust as first principles. That combination, not any single skill, separates a capable demonstration from a robot people can live with.

Humanoid service robot stands by a table in a modern living room. X Square Robot’s approach is that, in a real home, failure recovery matters more than raw success, because the home does not reset itself and it demands careful personalization, with safety and trust as first principles. X Square Robot

How do the open-source components fit into X Square Robot’s World Unified Model direction?

We see these releases as layers of the World Unified Model direction rather than isolated projects. Wall-OSS-0.5, the action model, asks whether an open vision-language-action model can gain directly measurable capability from large-scale pretraining, so it is the capability layer. WALL-WM, the world model, asks how a robot should understand change in the world, shifting from fixed windows to event-level modeling, so it is the representation layer. The data system supplies the interaction data that both of them learn from.

Together they form a loop in which models produce capability, world models organize understanding, and the open-source community drives reproduction and improvement. World Unified Model is the broader architecture those layers support, bringing vision, language, action, and physical prediction together.

We are releasing these pieces openly because embodied intelligence cannot be solved by one organization; it needs many embodiments, many real tasks, and broad feedback, and the long-term goal is a stack that keeps learning and ultimately moves robots from laboratory demonstrations toward reliable everyday use.

The foundational elements of AI architecture that IT leaders need to scale

With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces risk, leaving IT leaders to wonder which investments will prove valuable even six months into the future.

Returning to the foundational elements of AI architecture—the structural framework required for deploying and managing reliable, integrated AI systems at scale—allows technology leaders to make astute decisions today while supporting a future of AI agents that can retrieve information, make decisions, and execute complex workflows across systems.

Four elements of AI architecture you can count on

The following capabilities provide a stable compass on the path to production-ready deployment, regardless of how the underlying technology evolves.

1. Prepare data for AI at scale

Models are only as reliable as the data they can access, and poor data quality leads to AI hallucinations, bias, and unreliable outputs.

Most enterprises rely on legacy systems, inconsistent data structures, fragmented ownership, and incomplete datasets, making it difficult to scale AI effectively. Powerful as it is, AI itself cannot solve these underlying data problems.

As Adnan Adil, CIO of Elastic, explains: “The data is a durable part of AI architecture because without it, these models won’t run, won’t provide the right context, or won’t give the right level of services that we’re looking to implement.” Industry surveys consistently cite data quality as one of the greatest barriers to AI success. “The data quality has to be good; otherwise, the user loses confidence in the system,” says Adil.

An effective AI strategy begins with connecting data across the organization and ensuring it is organized, accurate, governed, and accessible in real time. These considerations are most effective when built into models and architecture from the start. Scalable data architecture allows AI systems to evolve alongside the business and connect reliably to the internal information needed to deliver meaningful value.

Gartner predicts that companies will abandon 60% of all AI projects through 2026 if they are not supported by AI-ready data. Avoiding that outcome includes clear data standards and ownership, clean and labeled data, and pipelines that support real-time retrieval.

2. Use context engineering to deliver the right data to every AI query

Context engineering ensures that the model draws on the most pertinent information for each query, selecting and organizing the data needed to produce accurate answers efficiently.

Effective context engineering shapes the inputs that guide AI reasoning and action. While prompt engineering focuses on how a request is worded, context engineering designs the entire information environment around the model: retrieving the right data and presenting it in a structured, machine-readable way. Many organizations are discovering that reliable AI depends as much on context quality as on the strength of the model.

Context engineering relies on a modernized, unified data foundation as well as retrieval and memory systems such as retrieval augmented generation (RAG) and vector databases. It also requires careful prioritization to determine what information matters most, what should be excluded, and when different types of information should be used. Feeding models too much context can dilute relevant details, increase costs, and slow response times.

“Minimum context, correct and current data, and machine-readable information are critical to effective context engineering,” Adil says.

3. Build AI governance and LLM observability in from the start

Strong governance and LLM observability help organizations maintain control over how AI systems use data, monitor system performance, and identify problems before they affect operations.

In the absence of clear controls around retrieval, workflows, and model usage, AI systems often process far more information than necessary. This inefficiency also drives up operating costs by requiring additional computing resources, often reflected in higher token consumption and API charges.

Governance also works in tandem with robust security. AI expands the attack surface, introducing risks such as prompt-based data leakage, model vulnerabilities, and adversarial inputs. Protecting sensitive information requires strong access controls, monitoring, and oversight.

Adil notes that essential controls — including those related to security, granular cost management, project controls, data security, and architecture—are frequently insufficient.

For governance systems to support transparent, compliant, trustworthy, and cost-effective AI, organizations cannot leave them as a layer to add later. Governance structures need to be embedded into architecture, workflows, and decision-making processes from the outset.

When governance is established from the start, it enables robust observability. Observability helps organizations understand how AI applications are performing in practice. Mechanisms for LLM observability and benchmarking allow teams to assess accuracy and utility over time, monitor adoption patterns, and adjust systems as conditions change. Observability also helps organizations gain trust by increasing visibility of model performance, behavior, and failure points.

Furthermore, observability is essential to get ROI of AI initiatives, as the benefits of it are often indirect and business value depends heavily on how systems are adopted and used. Real-time visibility into AI behavior allows organizations to measure performance against expectations, identify gaps between intent and reality, and continuously refine systems as requirements evolve.

In a 2026 report from Elastic, 85% of IT decision makers expect to enable LLM observability for their internal generative AI apps.

“Observability is actually huge. We can use observability data for cost control, decision-making, and engineering efficiency,” Adil says.

4. Keep humans in the loop

The thoughtful design, integration, and governance that maximize AI value demand specialized in-house expertise. Nearly 70% of respondents in Deloitte’s 2025 Tech Executive Survey report plan to grow teams in direct response to generative AI, a clear contrast to widely reported AI-related cuts. Adil agrees: “We think the people aspect is largely what’s going to make AI impactful going forward.”

As AI systems become more embedded in operations, organizations need people who can govern workflows, evaluate outputs, redesign processes, and adapt systems as conditions change. Evolution toward increasingly autonomous tools requires teams skilled in prompt engineering, orchestration, and change management. 

Talent adept at critical thinking and prepared to adapt with technology’s rapid advances will be in high demand. Although turnover brings in fresh thinking, it also presents high costs in system continuity, institutional understanding, and innovation. Human-centered strategy needs to be built into AI execution stages to ensure smooth implementation. 

As Adil says, “Many aspects of the stack are moving very, very fast, but institutional knowledge and the ability to adapt remain durable.

Thoughtful AI investment for future growth

As AI systems evolve from single-task assistants to increasingly autonomous agents, the organizations best positioned to benefit will be those that invest in the underlying systems, governance, and expertise that make AI reliable at scale.

Tech leaders who focus on these fundamentals can move effectively from experimentation to reliable, production-level deployment in the medium term, confident that these elements will remain relevant and adaptable amid constant advancements.

“We fundamentally believe that with these tools, velocity of work will get much faster,” Adil says. “We are really focused on how we can do work with these tools in ways we had not thought of before.”

Learn more about how Elastic is building an AI-first enterprise with these core foundational components.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Achieving operational excellence with AI

Frameworks like Lean Six Sigma and business process management (BPM) first gained traction because they promised clarity in the chaos—a structured way to bring order to messy, sprawling operations. Lean Six Sigma emphasized statistical rigor and quality control; BPM created end-to-end maps of how work should flow across departments. Both offered a repeatable way to embed habits of measurement, analysis, and accountability into day-to-day company culture.

But today, those time-tested playbooks are evolving as companies seek to embed AI into established process excellence methodologies. By some estimates, the market for AI-powered process optimization is projected to exceed $113 billion within the next decade. In one study, a full 88% of business leaders anticipated increasing investments into AI-infused process intelligence in the next 12 to 18 months.

Yet without the right foundations, many of those investments may not fully deliver on their potential. Companies that already operate with discipline have an edge. They can channel new tools into proven systems rather than bolting them onto shaky foundations. Organizations with mature process disciplines are also better positioned to translate AI ambition into real outcomes, as they are already accustomed to data-driven decision-making and process discipline—precisely the cultural foundation AI systems need to deliver value.

Simply put: AI can accelerate process excellence, but existing process excellence is what makes AI truly impactful. Technology and process are no longer separate levers, and only organizations that pull them together stand to realize the full value of both.

Download the full report.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Building the foundation for an autonomous enterprise

Artificial intelligence may have captured the public imagination through chatbots and image generators, but some of its most consequential use cases are unfolding far from consumer-facing tools. In industries where physical infrastructure, operational continuity, and safety are paramount, AI is becoming a core operating layer. With its sprawling industrial systems and constant stream of operational data, the energy sector offers a glimpse into what that future could look like.

At Woodside Energy, AI adoption did not begin with generative models or enterprise copilots. The company has spent years building predictive analytics, optimization systems, and machine learning tools across exploration, drilling, maintenance, and plant operations. “We’ve always had very large volumes of operational data coming from the equipment and the plants and the assets that we operate,” says the company’s vice president for digital Andrew Melouney. “Those have created really clear, quite high-value use cases for us.”

That long-term investment in infrastructure and governance is now enabling a broader shift toward agentic AI systems that can support complex industrial workflows. Rather than replace human operators, Woodside designs AI systems to augment expertise in high-stakes environments. A prime example is its “Startup Advisor,” an AI copilot that helps operators manage the complex process of starting liquefied natural gas (LNG) plants. “We’re really thinking about, how does it support the people in the organization in terms of empowering them to make better decisions, to make faster decisions,” Melouney explains.

The company’s approach reflects a wider evolution taking place across industrial AI: graduating from isolated experiments to enterprise-wide systems built on standardized platforms, governed data, and repeatable deployment patterns. That transition, Melouney argues, requires organizations to rethink both their technology stacks and how work itself gets done. “We’re not just bolting AI onto an existing process,” he says. “We’re deeply thinking about how that work needs to be reimagined.”

Melouney’s motto has become: “Think big, prototype small, and scale fast.”

As AI systems become more autonomous and interconnected, the companies poised to succeed may be those that spent years building the operational foundations beneath the hype.

“Our ambition is really for an autonomous enterprise, where we have agents with agency that are able to really deeply interact with our core workflows,” says Melouney.

This episode of Business Lab is produced in partnership with Infosys.

Full Transcript:

Megan Tatum: From MIT Technology Review, I’m Megan Tatum, and this is Business Lab, the show that helps business leaders make sense of new technologies coming out of the lab and into the marketplace.

This episode is produced in partnership with Infosys.

Now, when people think about artificial intelligence, they often picture chatbots or productivity tools, but some of the most sophisticated and high impact uses of AI are actually happening far from consumer apps, inside complex industrial environments where safety, reliability, and physical systems matter. The global energy sector is a prime example.

Companies like Woodside Energy, a global energy producer headquartered in Western Australia, have been applying AI for more than a decade now, from advanced analytics and operations, to remote decision support, to smarter maintenance, and energy efficiency across large scale assets. Today, Woodside is scaling that experience, embedding AI more deeply across its operations and the enterprise with a strong focus on governance, data quality, and human accountability.

Two words for you: technological fuel.

My guest today is Andrew Melouney, vice president for digital at Woodside Energy. Welcome, Andrew.

Andrew Melouney: Thanks, Megan. It’s great to be here.

Megan: Lovely to have you. Now, Andrew, as I said there, the energy sector has approached AI quite differently from technology or consumer businesses. Early value has emerged in operational and industrial environments, rather than consumer-facing generative AI tools. Why is that? And what differentiates the energy sector’s AI journey?

Andrew: Megan, I think it really comes down to the nature of the work we do. Energy operations and what Woodside does is very asset intensive, it’s very safety critical, and it’s highly physical. And when you think about how Woodside operates, we operate across the full value chain. We do exploration through to drilling and subsurface work, to project development, all the way through to operating assets, which are often operated in harsh and remote locations, and then global energy portfolio marketing and trading as well.

We’ve always had very large volumes of operational data coming from the equipment and the plants and the assets that we operate, and those have created really clear, quite high-value use cases for us. When you think about reliability, when you think about safety and efficiency, those are really critical things for a company like Woodside. We’ve been doing traditional AI for many years now. If you think about analytics, if you think about optimization, if you think about things like predictive models, those techniques we’ve been applying to our data sets and to our business since around 2015.

And more recently with the advent of generative AI, we’ve really found that we’ve got a pretty strong and awesome foundation to build on top of and to really solve problems in the service of improving the business. And again, whether that is keeping people safe, keeping the environments we operate in safe, or improving returns for the organization.

Megan: Fantastic. I mean you touched on it there, but how has this reality shaped your own AI strategy at Woodside? Where did you start, and where did the technology prove most impactful in those early days?

Andrew: Well, like I said, we’ve had a very long journey, in terms of understanding our operational data, recognizing the value of it, and collecting it at scale so that we can use it. And we’ve been very deliberate in that approach, Megan. We’ve really thought about where the value is and where the risks were manageable. And we’ve started looking at, in today’s world from an agentic AI perspective, we’ve started looking at the problems that were solved with traditional AI and machine learning and data science in the past. And we’ve started to think about, where can we then layer agentic AI over the top to provide an even better outcome?

For our asset intensive industry and organization, we’re looking at areas such as maintenance optimization. We’re looking at areas such as, how do we ensure our LNG plants start up reliably, consistently, and safely? And we’re considering really our frontline workforce and making sure that we’re giving people on the frontline the tools required to do their jobs. When we think about AI, we’re really thinking about, how does it support the people in the organization in terms of empowering them to make better decisions, to make faster decisions? I think over time, this has just evolved from what has been traditional analytics to now artificial intelligence and generative AI. And we’ve learned along the way that the technology is important, but it’s about aligning people, processes, and the technology together.

We’ve spent a long time not only in collecting the data and having a well-curated data set that we can build on top of, but we’ve also spent a lot of time teaching people how to work in agile ways, how to do design thinking, how to problem solve, and how to really make sure that the technology that, say, my team can bring to bear to the organization is adopted effectively and purposefully. And I think once we had that solid foundation in place from a technology perspective, from a data perspective, once we got strong trust built between our digital teams and the organization, we really saw quite a material uptick and the scaling of technology occur more broadly across the enterprise.

Megan: Fantastic. That people piece so important, isn’t it? It’s just a tool, technology, that needs to be in the right hands. And you touched on data there; industrial AI obviously depends on vast amounts of data. Can you walk us through how you’ve approached data at Woodside in a little more detail? How it’s structured and governed, and how tools like maintenance intelligence as well fit into that.

Andrew: Well, data is really foundational and fundamental to everything we do, particularly from a technology perspective. It gives us the ability to innovate at pace when we are building over the top of a strong foundation. As I said before, we’ve had the benefit of a long-term investment in our underlying operational data. I think the way we think about data is that it’s an asset for us.

And when you think about operating a facility where you’ve got sensors everywhere, you’ve got data streaming in real time, you’ve got operators needing to make decisions in real time, we have consciously made a decision over many, many years to invest in that enterprise scale data platform to make sure that it’s secure. We’ve got well-structured data assets, and we’ve got strong governance over the top of that data so that when it is used, when it’s built in a data science application or an AI agent, that we’ve got a level of trust in it that it’s going to be used responsibly. And that when it’s used, it can be trusted to give the outcome that we expect.

We have developed platforms that continuously ingest really high frequency data from the assets and from our enterprise systems. Once we’ve been able to develop solutions on top of that, parts of the business that might own the systems that collect that data, they see the value in it.

When you look at something like maintenance intelligence is a really good example of how we’ve been able to take something that we’ve been working on for a long time. Woodside does a lot of maintenance, it’s a very important part of our business, and it occurs across all of our operating assets. But we have been looking at how we do predictive analytics and predictive maintenance for a long time across that data set that we own. And something like maintenance intelligence is a solution that gives us the ability to optimize how we do that maintenance. And what it does is it analyzes historical maintenance records, alongside the performance of the equipment. And again, by having that data set well-governed and in one place, we get the ability to correlate different data sets, such as maintenance records out of SAP, alongside say equipment and performance coming from our time series data lake.

And when we build over the top of that, something like maintenance intelligence gives us the opportunity to recommend to the assets what the optimal timing for maintenance activities might be, and really give what is quite a simple aim, which is do the right work at the right time. And with something like maintenance intelligence, we have seen the opportunity, and we have the opportunity to reduce maintenance hours by up to 15% over five years on one of the assets that we’ve piloted this on. And as we’ve built out that underlying analytical model, we’re now able to put agentic AI over the top of that and provide better insights and optimize that solution more.

It really comes down to providing our asset teams and our operational teams with the right decision support capability that ensures they’re still accountable to make the decision and to ensure the right work is being done, but we are giving them the best possible opportunity to use their judgment and experience with the data that we provide to make the right decision.

Megan: Sounds like a really impactful change. Last year also marked a milestone in moving from early AI learnings to scale, using AI more deliberately as a force multiplier. What transition were you trying to make and how did you approach it?

Andrew: Well, Megan, we’ve had a philosophy for a long time in Woodside from an innovation perspective, where we really want to think big, we want to prototype small, and we want to scale fast. We want to find big opportunities that we can go after, but we want to ensure that we look at how we deploy those on a small scale first, and then provide the right learning and insight that then can scale it everywhere. Something like maintenance intelligence is a good example of that, or our Startup Advisor, where we know that we’ve got multiple plants that we need to start up. We know that we’ve got multiple assets that need to do maintenance, so we have a big, bold ambition about how we can improve and optimize that. We start with a small prototype; it might be one subsystem, it might be just a part of an asset, and then we scale it out, we learn, and we scale faster.

I think from an AI learning perspective, one of the key things we’ve learned is really the transition from moving from isolated AI solutions to a more coordinated enterprise-wide capability. If you look back maybe 18 months, two years, in our generative AI journey, we rarely started by deploying AI as broadly as we could in the organization from a personal productivity perspective. And probably being quite open in terms of the problems that we will solve, the business problems that we’ll solve with AI. That had a lot of benefits for us in terms of allowing our organization to get to know AI, get to know the capabilities, to build the trust in it.

What we’ve learned though is that we’ve needed to pivot from that to being a little bit tighter in terms of where we are going to invest our time and resources and more higher value solutions. How do we then enable and empower the rest of the organization so that they can actually effectively problem solve with technology in their domain or in their personal productivity without having to come to a central team?

When we think about that, think big, prototype small, scale fast, has been something really important for us. The transition from a more broader approach to use case development and solution development to now a narrower focus on the high value priorities. We’ve seen that paying dividends to us and allowing us to go after solutions and opportunities, things like Startup Advisor.

And so our Startup Advisor is a agentic AI solution that really aims to optimize and empower and better support our operators that sit in front of a panel and have to start up LNG plants, which are incredibly technical facilities and require really specialist skills to start up. And so our Startup Advisor is almost like a copilot that sits alongside those operators, and it gives them the ability to be able to play back previous startups. It gives them the ability to look at how the current startup is progressing, and it provides them better insights to optimize how they start up that facility. And again, starting up an LNG facility is incredibly complex.

Megan: I can imagine.

Andrew: When we think about opportunities like Startup Advisor, again, it goes back to that think big, prototype small, and scale fast. We started with a very bold vision of, how do we start up all of our LNG plants in a much more structured and optimized fashion? How do we better support our panel operators? How do we make, say, a more junior panel operator have a copilot that can help them almost like an experienced panel operator sitting next to them? And when we think about that vision and the ability then to prototype on a small scale and then scale fast, I think it’s been really successful for us.

As we scale, we’ve just naturally expanded into more agent-based solutions. Today, we’ve got around 50 AI agents in production, supporting both our operating assets and our enterprise workflows. These tools have been proven in live environments, and we have really seen the benefit of being able to shift from point solutions that maybe solve small scale problems in specific areas, to AI and agentic solutions with agency that can really work across our workflows.

We’re able to do this because we’ve standardized on the platform that we build on and we’ve got repeatable patterns. That’s been another really important learning for us, is that we don’t want to build 50 solutions in 50 different ways. We really want to be empowering our organization and our technical teams and the users of our solutions to roll them out quickly, to roll them out safely, and to do it in a patternized and platform manner.

But the last point I’ll make, Megan, from a learning perspective is that we’ve really understood that a strong governance around how AI is deployed and developed is critical for us, and it’s critical for us to go fast as well. The traditional ways of governing how we roll out different solutions or digital systems isn’t going to scale to the breadth that we need when we are thinking about AI. Being able to have a clear philosophy around how we innovate, transitioning from isolated solutions to that enterprise-wide capability, and making sure that we’ve got strong platforms with strong patterns and clear governance are the three really critical things that we’ve learned.

Megan: Such important pillars, all of them. And you’ve been working with Infosys on this journey. How has that partnership helped accelerate scaling and embedding AI across the business?

Andrew: Well, Infosys is our managed service provider, and so they play a really critical role in the operations of our core business. One of the things that I like to say is that our license to innovate is based on our license to operate. And so, for my team to be able to turn up to an operating asset or a corporate function and have the trust that’s needed to be able to innovate and reimagine and redesign how work gets done, to be able to do that, we need to make sure that our core platforms, our core systems, our applications are running really reliably, safely, and consistently every day. Having an experienced partner like Infosys looking after those core operations in partnership with our internal teams is really, really important to us.

As we move from pilots to enterprise-wide deployment, the ability to partner with someone like Infosys also gives us the ability to scale. And so being from Perth and Western Australia, while we’ve got a really strong local team in Western Australia, and we’ve also got a very strong team in some of our other operating locations, like everyone, we’re struggling to find people that can fill AI roles. Being able to partner with Infosys and have a number of different operating models at our disposal becomes really important for us. Having co-mingled teams where they are staff, they are Infosys staff, Woodside staff, and some of our other partners, really just brings diversity of thought and experience to how we solve problems.

Fundamentally, the partnership has allowed us to operate and innovate with more confidence. While Woodside always retains ownership of the strategy and where we’re going and the governance and my teams remain accountable for the outcomes, we can’t do what we do without strong partnerships like the one we have with Infosys.

Megan: Fantastic. And as AI adoption scales, you mentioned yourself, governance becomes increasingly important. How challenging has that been, and what guardrails have you put in place at Woodside?

Andrew: So, Megan, governance is really important to us, and we operate in a well-regulated environment. That means we’ve got to make really deliberate and well-reasoned decisions when we’re thinking about how we deploy technology into our organization, whether it’s artificial intelligence or anything else, for that matter. And so, governance is really central to how we approach the execution of our AI strategy at Woodside.

We’ve got maybe two or three really key things that we’ve put in place. The first one is just making sure that every AI use case goes through a structured assessment, and that’s making sure it meets our privacy controls, our cyber controls. We’re also asking the question, not just, could we do this, but should we do this? We’ve really got to bring together safety, ethics, transparency, accountability, and make sure that we make an informed decision. When an AI solution is going through that structured assessment, if there are concerns about how we might use that solution, it then goes to an AI council that’s made up of senior leaders across the organization. That council and that group really oversee some of the prioritization and risk management. That’s where we can have really strong, robust debates around, again, could we do something, should we do it, and how do we mitigate any of the risks that we might introduce here?

I think the last one, Megan, is really around lifecycle management. When you start thinking about, we’ve got 50 at the moment, but if we had 500 agents working in our organization, really amplifying the experience and the decision-making and the value creation of our staff, we really want to have an ability to manage the lifecycle of how those agents operate. We want to know, how many people are using them? What’s the efficacy and the outcome? Is there model drift? Do we need to retune or retrain? I think that’s an area where many organizations, including Woodside, are still leaning into and still figuring out the best way to do this. We can do it quite easily with 50 agents, but 500, 5,000, 50,000 becomes an opportunity for us. Again, thinking about how we partner with others, solving problems like that really present an opportunity to co-create and to co-solve with some of our partners, like with Infosys.

Megan: Fantastic. Just to close, what’s your long-term vision for AI at Woodside? How do you see this evolving over the years ahead, and what could it unlock for the sector in your view?

Andrew: So Megan, I think our ambition is really for an autonomous enterprise, where we have agents with agency that are able to really deeply interact with our core workflows. The outcome that we want to get from that is to protect our people, to protect the environments we operate in, and to be able to provide energy at a lower cost to the world. When we think about that ambition, we can really see that being applied to almost all of the areas that Woodside work in. Whether that’s from exploration through to project developments, through to operations or marketing, the scale of the opportunity in front of us and the ability for us to really change the way that work flows through the organization is really exciting.

For us, there’s three things that we have to get right in terms of being able to execute on that ambition. The first one is really thinking about how the work gets done in the organization so that we’re not just bolting AI onto an existing process, but we’re deeply thinking about how that work needs to be reimagined. We’ve also got to think about how we enable our workforce to work differently. Providing them with the skills and the tools and the ability to really harness the power of the technology that we provide.

Secondly, we’ve got to continue to move from and restrain ourselves from deploying point solutions that solve very narrow problems, to having more connected, agentic systems of systems that can interact with each other. To do that, and if we do that successfully, that’s where we really get the high value unlock from agents being able to interact with workflows and really change how the work gets done.

And lastly, Megan, it’s about how we must continue our philosophy of thinking big, prototyping small, and scaling fast.

Megan: Which is a fantastic lens to which to make all these decisions. Thank you so much, Andrew. That was Andrew Melouney, vice president for digital at Woodside Energy, whom I spoke with from Brighton in England.

That’s it for this episode of Business Lab. I’m your host, Megan Tatum. I’m a contributing editor and host for Insights, the custom publishing division of MIT Technology Review. We were founded in 1899 at the Massachusetts Institute of Technology, and you can find us in print, on the web, and at events each year around the world. For more information about us and the show, please check out our website at technologyreview.com.

This show is available wherever you get your podcasts. And if you enjoyed this episode, we hope you’ll take a moment to rate and review us. Business Lab is a production of MIT Technology Review, and this episode was produced by Giro Studios. Thanks ever so much for listening. Goodbye.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Agriculture is ready for AI, but its data isn’t

Artificial intelligence is transforming what is possible in agriculture, but industry leaders should be wary of investing in AI without first laying the groundwork. 

The use cases are promising, especially for an industry navigating volatile fertilizer costs, unpredictable weather, and margins that leave little room for error. Research shows AI-enabled predictive models can improve crop yield by 26%, reduce water use by 41%, and cut chemical usage by 33%. 

However, what AI vendors usually won’t tell you is that these solutions are only effective if you have a clean, solid data foundation. However, at Reltio, we have experience in this area, including leading technology strategy at a major agricultural distributor and building a data platform used by enterprises worldwide–we’ve seen it first hand.

What AI vendors won’t tell you 

Vendor conversations in agriculture tend to follow a familiar pattern. The pitch leads with grand promises around using AI to monitor crop health in real time, optimize irrigation, and squeeze more yield from every acre. 

The promise is compelling, but what rarely comes up is the question of whether the data foundation underneath those promises is accurate and complete. If not, there is a real and significant risk that AI will generate misleading outputs that seem authoritative but inspire action that is, at best, counterproductive. 

For instance, a yield prediction model fed inconsistent historical data will generate imprecise forecasts. Similarly, a precision irrigation system drawing on fragmented sensor data will make watering decisions that waste resources instead of saving them. 

In each case, the AI is failing because the data it was trained on was not sufficient to produce trustworthy outputs. In agriculture, every AI hallucination is a liability, and the likelihood of error is high.

Why agriculture is a uniquely challenging test case

The data landscape across a modern agricultural operation or a large distributor serving thousands of growers is extraordinarily complex.

Modern farming environments make extensive use of IoT devices and machinery. Irrigation systems are automated, tractors navigate fields autonomously, and drones capture field imagery at scale. 

However, machine data is disparate by nature. Add in external sources, including weather feeds, U.S. Department of Agriculture data, and third-party market information, and the question of how you bring all of it together into something coherent becomes a significant undertaking. 

Agricultural AI also needs to understand more than just customer attributes; it needs to understand the land: GPS coordinates, farm boundaries, field blocks, and soil variation across a single property. Where do you apply fertilizer, and at what rate, and in which specific area of the farm? Not all parts of a field are the same, and an AI system that treats them as if they are will produce recommendations that are at best imprecise and at worst damaging.

There is also a compliance dimension due to the chemicals and the responsibility involved. Operational AI in agriculture needs significantly more checks and governance than it might in a lower-stakes environment. When a flawed recommendation gets acted upon in the field, the consequences can be severe. 

What data readiness means in practice 

Data readiness is the difference between AI delivering on its promise vs. a “garbage in, garbage out” scenario. Fundamentally, being ready for AI means having a data model that accurately reflects how the business operates. 

For a company like Wilbur-Ellis, a 104-year-old, family-owned agricultural distributor, that means understanding who your customers are, which fields they farm, which inputs they need, which suppliers those inputs come from, what they paid last season, and how all of that connects to margin. That information needs to be current, consistent, and accessible across the organization, rather than locked in separate systems that were never designed to talk to each other.

Similarly, for farming operations themselves, data readiness means having a reliable, connected picture of what is happening across every field: soil health records, input application histories, yield data from previous seasons, equipment performance, and real-time sensor readings from irrigation systems.

Governance matters just as much as structure. Prices change, relationships evolve, and suppliers come and go. An AI system drawing on data that was accurate six months ago but has not been maintained will make recommendations based on a version of the business that no longer exists. 

Building the foundation that makes AI trustworthy

The good news is that the path to data readiness is feasible. It starts with a strong data model: a single, governed source of truth that connects customers, suppliers, products, pricing, orders, and margins in a way that reflects how the organization operates. 

From there, it requires data pipelines fast enough to deliver insights when decisions need to be made, governance frameworks that keep that data trustworthy over time, and security controls that ensure sensitive commercial information is accessible to the right people under the right conditions.

This is precisely the challenge that Reltio, an SAP company, was built to solve. Reltio enables companies to unify their fragmented data so AI agents and systems can operate from a complete picture of the business. Reltio builds a trusted system of context, known as the context intelligence layer, that brings all entities, relationships, rules together under one roof and makes business data easy to access and interpret.

For Wilbur-Ellis, building that trustworthy data foundation has meant being able to ask more complex questions and trust the answers, which is the precondition for any AI system to be genuinely useful.

How agriculture can drive real value from AI

The question worth asking before the next AI conversation is not whether the use case is promising. It almost certainly is. The question is whether the underlying data foundation is strong enough to make the output trustworthy. 

Agriculture has always required its leaders to make high-stakes decisions under uncertainty, and AI offers the genuine prospect of making those decisions faster and better informed. That prospect is only achievable for organizations that have done the foundational work first, and the businesses that will get the most from AI are the ones investing in that foundation now.

This content was produced by Reltio. It was not written by MIT Technology Review’s editorial staff.

Agent confidence on the technical frontier

Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressure to prove ROI mounts, executives and technology leaders are looking to agentic AI to drive the measurable financial outcomes their businesses seek.

A prime opportunity for AI agents exists in the tech function, where IT infrastructure costs are projected to grow two to three times by 2030, even as budgets remain unchanged, according to McKinsey. And in the last 18 months, tech teams—the engineers, developers, architects, and other practitioners who are building, deploying, and continually improving their organizations’ infrastructure and applications—are clearly putting agents to work.

The ultimate promise of agents is not only to automate tasks but to manage and coordinate entire workflows, pursuing business goals in a way that allows humans and agents to work together. Given the risks involved in automated decision-making, teams cannot delegate the work that agents do without confidence that they are fully capable of performing the task and that it will do so in a safe, reliable, and secure manner.

Among technology experts, our research shows that teams are exceedingly confident about using agentic AI across a significant amount of AI, data, and cloud tasks.

Where agent readiness drops is largely due to a lack of business context being supplied to agentic systems. The more complex the task, the more reasoning capability an agent requires and the greater its need for business context. Such context-generation capabilities for agents are still at an early stage of development, especially in situations where enterprise data is difficult to wrangle and connect into the agent lifecycle at the speed and quality in which developers and executives need it. Human oversight is a key factor of success in deploying agentic AI.

Knowing that tech teams are in a pivotal position to lead this transformation, the experts we interviewed expect agent confidence to accelerate as experience with agents deepens and business environments mature. “As we design agents to operate within the same operational boundaries, identity systems, and governance models that teams already use, they start to behave more like the systems organizations already trust,” says Jeremy Winter, corporate vice president and chief product officer at Microsoft Azure Platform.

This report, based on a survey of 300 global technology experts, ranks 101 tasks across AI, data, and cloud workflows based on respondents’ confidence in agents acting on their behalf. It also examines how technology teams view the opportunities and challenges related to agentic AI, along with the potential for the technology to enhance their careers.

Key findings from the report include:

Confidence in agents is surging for measurable tasks and growing in areas of complex judgment. Technology experts overwhelmingly believe agents help with everyday work including streamlining processes, improving performance, and reducing repetitive tasks. Confidence is highest for processes like generating reports and boilerplate code, and there is clear opportunity where tasks involve multistep workflows and advanced reasoning to make decisions.

Data workflows are the breakthrough domain. Tech teams trust agents most where structure can provide a reliable foundation for decisions. This includes areas such as data quality monitoring, visualization anomaly detection, real-time data stream monitoring, and data profiling. This is where domain experts closest to the point of data generation can provide context to allow agents to act and deliver trusted outcomes.

Download the full report.

Read the Microsoft Cloud blog by Amanda Silver, corporate vice president of Microsoft 365 Core and Work IQ, which underscores the importance of keeping humans in the loop and how systems thinking advances careers. And for a deeper dive into data workflows as a breakthrough use case for agents, check out the Fabric blog to hear from Kim Manis, corporate vice president of Product for Microsoft Fabric.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.




Repositioning retail for the AI era

Artificial intelligence is rapidly reshaping retail, but not in the ways consumers might immediately notice. The biggest transformation may not be flashy virtual try-ons or chatbot shopping assistants, but in how decisions are made behind the scenes: how products surface in search results, how inventory moves through supply chains, how engineers ship code faster, and how retailers respond to customer behavior in real time. As legacy retailers navigate a fragmented and hyper-competitive landscape, AI is becoming an operating philosophy.


At Macy’s, that philosophy is more often defined by what senior director of engineering Murali Murugan describes as an “AI-first” approach. “AI first isn’t about adding intelligence on top,” Murugan says. “It’s about redesigning how decisions happen so the business moves faster and every experience feels more relevant by default.” Rather than layering AI onto existing workflows, Macy’s is embedding intelligence directly into systems that include personalization, search, operational planning, and software development itself.

The company’s strategy is reflective of a larger shift taking place across retail: moving from isolated AI pilots toward integrated systems designed to compress, as Murugan puts it, “the gap between the signal and the action.” Early efforts focused on narrow, high-impact use cases like search recommendations and customer engagement, where measurable gains in conversion and reduced friction quickly built internal momentum. “Once we established the quick wins, scaling was a business decision, not a technology debate anymore,” he says.

That momentum is now extending into conversational commerce through tools like Ask Macy’s, an AI-powered shopping assistant designed to act more like a personal stylist than a traditional search bar. Whether for a prom, a vacation, or a last-minute event, customers can describe what they need conversationally and receive curated recommendations informed by past purchases, preferences, and context.

Still, the company sees AI as more of an invisible layer augmenting human judgment than a replacement for it. The long-term vision is retail that feels increasingly seamless, adaptive, and personalized, powered by systems customers may never even notice are there.

“The real transformation in this all comes from continuous improvement,” Murugan says. “It’s about learning from the mistakes, quickly adapting to the newer technology standards that are coming into play, timing, and execution which compound into a meaningfully better customer experience.” 

This webcast is produced in partnership with Infosys.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

❌