Normal view

The path to artificial superintelligence

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy.

Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate patient care without a human making the decisions.

“The intelligence is already there. What is missing is the connective tissue that turns four strangers into one team,” explains Vijoy Pandey, senior vice president and general manager of Outshift by Cisco.

This “connective tissue” comes from adding a semantic layer—what Outshift calls the “Internet of Cognition”—that enables agents across domains to work together and, critically, “think” together through shared intent, context, and reasoning.

This semantic layer relies on a connectivity layer beneath it called the “Internet of Agents,” which allows autonomous agents to discover one another, prove identity, and exchange messages across domains.

When used together, they enable “the next step on the road to distributed artificial superintelligence,” says Pandey.

From solo silicon savants to the ‘Internet of Cognition’

For years, the AI industry has been focused on growth. Scaling vertically has led to bigger models, trained on more data with more compute. This has produced the reasoning capabilities that can be like a “brain” for AI agents, which can perceive, reason, and act in digital environments.

While vertical scaling can produce more capable agents perpetually, to enable agentic problem solving across different systems, companies, and platforms the next axis of scale must be horizontal, says Pandey.

Multi-agent systems are already being explored in areas like software engineering, drug discovery, and scientific simulations, but their performances so far have been underwhelming. One study finds a failure rate of between 41% and around 87% when evaluating seven open-source multi-agent systems.

“Connected agents handle coordinated action well; taking a task whose shape they have seen, divided and passed around,” Pandey explains. “What they cannot do is hold a goal in common and reason toward something none of them was trained to solve.”

“The gap is architectural, not a prompting problem,” Pandey adds. “Without the right coordination layer, naive multi-agent setups can perform worse than a single agent. The step change is that team of agents converging on its own, on a new problem, with no human stitching the seams.”

To reach this goal, Pandey says Outshift has built a connectivity layer called AGNTCY, an open-source project now under the Linux Foundation. AGNTCY allows agents across different systems, companies, and platforms to find each other, prove identity, and exchange messages through open, standardized protocols.

And, as Pandey explains, this allows the Internet of Cognition thesis to take a step further. It creates a semantic layer that allows agents to align goals (share intent), pool institutional knowledge and compound memory (share context), and make collective trade-offs (share reasoning).

Pandey likens this progression to that of humans: “For hundreds of thousands of years humans got individually smarter, and the gains died with each person who made them,” he explains. “Around 70,000 years ago that changed, when humans learned to share intent, build cumulative knowledge, and reason collectively. That is when scattered individuals became civilization.

“Agents are at the same threshold. We have built the silicon geniuses and given them agency. What they lack is the layer that let humans go collective,” he says.

First steps to distributed superintelligence

Enabling agents to work collectively rests on three pillars in the tech stack:

Shared intent through cognition state protocols: Cognition state protocols are the semantic handshake that allow agents to agree on a goal before they act and then negotiate toward it. Outshift has created an open-source coordination layer called Mycelium, which organizations can clone and use against their own agents.

“We found that unstructured groups reached a decision about a third of the time across 14 scenarios,” says Pandey, speaking about internal testing. “A coordination protocol that makes agents declare a goal, surface missing information, and resolve conflicts before acting raised that to 93%.”

Shared context through cognition fabric: A cognition fabric is a shared institutional memory and communication mesh that allows agent insight to compound over time rather than resetting each session. This policy-governed context layer solves the problem of “organizational amnesia,” says Pandey, by ensuring the baseline intelligence of the systems only ever goes up.

Shared reasoning through cognitive amplifiers and guardrail technologies: Two kinds of cognition engine can be used together to enable shared reasoning. Cognitive amplifiers speed up shared reasoning and modeling, and guardrail technologies (GATs) create security, cost, and compliance frameworks. Humans are active contributors to this layer, making judgment calls the system routes to them (rather than reviewing outputs after the fact).

Cognition sharing in multi-agent systems can create new risks, including unintended delegations, malicious prompt injections or memory poisoning, or over-privileged agents with access to permissions and data far beyond what their tasks require. Environment-specific controls are therefore needed to protect against unintended actions or consequences.

“Agents have human-like attributes but operate at machine speed and scale,” says Pandey. “Everything we built for twenty years—access control, identity, compliance—was built for humans or machines, not both.”

Continuous Agent Semantic Authorization (CASA)—an open-source reference implementation developed by Outshift—is a GAT that works to ensure agent actions remain securely aligned with the user’s original goal through a process of continuous authorization. It does this by reading what the agent is trying to accomplish then checking each tool request against that task.

In the case of a healthcare system, for example, an agent told to summarize a patient record may start by querying a whole database. This could lead to CASA denying the call, because the request no longer matches the task it was authorized for.

“Today’s controls are scoped to a role or a session not to the task so an agent granted a tool can use it for anything,” explains Pandey. “Roughly 90% of the time, an agent has no way to confirm it is even cleared for the job it was handed.”

Experimentation for cross-domain innovation

When horizontally scaling intelligence in the enterprise, businesses should begin by experimenting with one cross-functional workflow that spans three or four teams and currently needs a human authorizing the handoffs, Pandey advises.

“Stand it up as a small multi-agent system on open, interoperable infrastructure, with a measurable baseline,” he says. “Keep building bigger models, add the horizontal axis on top of them, and change what you measure. Track where one agent’s insight made another agent better—that is the signal the horizontal axis is working.”

By starting to experiment now with intent, context, and reasoning layers, organizations can get ahead of the curve. “The problems are open, and the infrastructure is still being written,” says Pandey. “This is the moment to build it.”

For more information on the Internet of Cognition, visit Outshift.com.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.




Closing the data loop in AI-driven drug discovery

Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage.

Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years—a phenomenon known as Eroom’s Law. Today, bringing a new drug to market takes an average of 10-15 years and costs anywhere from $1 billion to $2.5 billion, with failure rates upward of 90%.

AI has become the pharmaceutical industry’s biggest bet on bringing success rates up and timelines down. The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in development.

“The main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial,” says Paul Belcher, director of protein research strategy at global life sciences company Cytiva. “AI is one approach that drug companies hope will not only save time and compress timelines, but enable better quality candidates to reach the clinic.”

Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in lab systems.

AI brings efficiency to the lab

One of the most promising early-stage applications of AI in drug discovery is in hit identification. This involves screening libraries of molecular entities against a disease-related target, such as a protein, to find molecules that bind to it. A successful hit gives researchers a starting point for further testing and refinement, with the aim of eventually developing a viable drug.

Belcher has seen a shift from empirical screening to predictive design: Instead of physically screening libraries, drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing anything to research and development (R&D).

This means companies are no longer limited by how much they can physically screen to identify starting points. “AI does away with that,” says Belcher. “And it can help eliminate low-quality candidates before you have to physically test them, saving time and resources.”

What AI can’t do yet is reliably predict kinetics or developability of new compounds, says Belcher. This means every AI-generated candidate still needs to be validated in the lab.

Traditional screening workflows were built to identify hits at scale, not to profile large numbers of complex candidates in detail. This is placing more pressure on lab teams, who now have to test, characterize, and purify a growing volume of more diverse, AI-generated compounds.

“The current techniques used in hit identification can screen hundreds of thousands, sometimes millions of compounds, using binary or threshold-based techniques producing low-fidelity data—yes-or-no responses,” Belcher explains. “AI can increase the number of hits you get and potentially give you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits.”

Models need complete, quality data

As AI has accelerated demand for data-rich lab systems, it has also highlighted a fundamental need for better, more complete data.

Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. Because models have access to the same data, they all reach similar conclusions, with diminishing returns over time. Additionally, the datasets weren’t built with AI in mind, meaning they lack the structure, labeling, and diversity needed to keep models accurate and free of bias.

Publication bias reinforces the problem. “Most publicly available datasets and scientific publications focus exclusively on positive results,” says Belcher. “No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable.”

The data Belcher believes would markedly improve models—the failed experiments, the compounds that don’t bind—remains frustratingly difficult to come by. “We often joke that there should be a journal of negative data,” he says. “It’s often buried in lab notebooks, and it’s never used to inform or guide future research.”

This lack of negative data creates a fundamental problem: Without access to a broad range of data, models can’t be adequately trained to avoid bias. “In all machine learning applications, the model’s performance relies heavily on the quality and scope of the training data,” notes Belcher.

Fabrication has also become much easier with AI, compounding concerns around data integrity. Take Western blots, for example. These are part of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. This was back in 2016, before generative AI made fabrication trivial.

“Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences,” says Belcher. “There needs to be more tools to verify that data is not manipulated.”

Some vendors are starting to tackle this challenge. Belcher points to solutions like Cytiva’s Image Integrity Checker, for instance, which uses secure hash algorithms—the same technology used in blockchain—to detect whether scientific images have been tampered with. “We’re starting to see a lot of interest from publishing houses that want to adopt this as standard because it’s a quick way to ensure that what gets published in the literature is genuine,” he adds.

Autonomous labs could accelerate breakthroughs

Belcher describes the future state of drug discovery as fully autonomous labs that run with minimal human intervention. Foundational to this vision is consistency in data and infrastructure.

These AI-driven dark labs, or labs-in-the-loop, operate around the clock. They cycle through prediction, testing, and optimization, and then feed results back into AI models to guide the next round of experiments. This can improve the success rates of drug candidates entering clinical trials, says Belcher. Better starting points, combined with more rounds of optimization, should result in better candidates with fewer liabilities reaching the clinic.

But automating a lab depends heavily on integration. That means interoperable systems, highly structured and comprehensive datasets, and information flowing easily in and out. Most labs aren’t there yet. “Today, a lot of the instruments in labs are standalone,” Belcher notes. “You can have the best technology in the world, but if it’s a closed ecosystem—if the user can’t get the data out—it doesn’t do any good.”

An integrated infrastructure can enable labs to generate FAIR (findable, accessible, interoperable, and reusable) data at scale. This would not only inform individual lab reports, but could also train subsequent generations of AI models, effectively closing the loop between the computational, AI-driven dry lab and the physical wet lab.

“Our goal is to help scientists and researchers accelerate their breakthroughs and make that future state of autonomous labs a real possibility,” says Belcher. “We want to help them generate reliable data, simplify workflows in discovery, and hopefully enable what they’re working on to become tomorrow’s life-changing therapies, faster and with greater confidence.”

On costs and what comes next

AI-driven drug discovery is still in its early days. Notably, no drug discovered primarily through AI-driven design has yet received full FDA approval—although Belcher expects that to change in the next two to three years.

How big of an impact could AI eventually have on drug discovery? “The holy grail would be full in silico prediction of efficacy and toxicity, eliminating the need for the vast majority of physical wet lab work,” says Belcher. But there are many barriers to this beyond the maturity of the models, including regulatory hurdles and cost challenges.

A Stanford study found that the cost of training frontier AI models has more than doubled every year since 2016, adding more financial pressure to a sector already defined by exceptionally high R&D spend.

Belcher acknowledges the tension, but remains optimistic about what’s ahead. “I think we’ll get to a point where there’s a balance between AI and wet work, from a cost perspective and a risk perspective,” he says. “As long as the cost of compute doesn’t ever outweigh the cost of clinical development, I think AI is going to be an advantage.”

Learn more about how Cytiva is using faster discovery to reshape protein purification workflows.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

Building the enterprise environment for agentic AI

For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the ability to predictably plan and scale agents.

To better understand some of these dependencies, Intel performed thousands of agentic AI workload experiments. Our initial findings create and support five practical lessons for enterprise leaders:

  1. Agentic AI is a larger systems problem, not just one of inference.
  2. The majority of existing agentic AI harnesses are limited and do not measure overall system performance.
  3. Plan capacity is done using agents per virtual CPU (vCPU) density, not agent count.
  4. Monitor agent task latency, not just average CPU utilization.
  5. Default to scale-out for systems hosting agents. Reserve scale-up for workloads with heavier per-agent compute or architectural constraints.

Beyond inference: Agents as workflow automation

Agentic AI is more than LLM inference. Its enterprise value depends on the full system, task orchestration, data access, tool execution, latency management, governance, and scalable infrastructure. An agent is a goal-driven automated enterprise workflow process: It plans a multi-step task, calls tools, reads results, and retries when something fails. Enterprise agents are therefore not just an inference problem; they are a systems problem.

Defining what good looks like

Most agentic AI metrics focus on evaluating the LLM used. Platform teams also need to know how long the tasks take, how many agents a fleet can support, what users experience at the end of the execution process, and how costs change as more agents work simultaneously.

A more useful enterprise view looks at six metrics:

  1. Task success rate
  2. Cost per task
  3. Time per task
  4. Task throughput
  5. Agent density (agents per vCPU)
  6. Latency

Together, these answer the questions enterprise AI operators care about: Is the system performing as expected? How many agents can the system sustain? How should it scale to support more agents?

Building on solid foundations

To gain a deeper insight into agentic AI workload performance, Intel extended Terminal-Bench, an open source benchmarking harness for evaluating AI agents with profiling, telemetry, and replay capabilities. This made it possible to understand where the agents spent time beyond LLM inference.

The benchmark extension used a deterministic record-replay of LLM responses to separate agent performance from LLM variability. LLM responses were recorded once and replayed identically across runs, reducing run-to-run variance and creating a more reliable basis for comparison.

The Terminal-Bench task mix used was intentionally broad. It included compilation, testing, database operations, Boolean logic, interpretation, ray tracing, compression, linear algebra, video transcoding, and machine learning training. That wide variety made the findings more relevant to real enterprise environments.

Agentic AI in three dimensions

Deploying agentic AI should be approached in three phases:

Plan in terms of agent density, not agent count: The first sizing rule is to normalize agent count by available compute. Agent density, measured as agents per vCPU, is the leading signal for saturation. For example, 10 agents on an 8-vCPU system and 20 agents on a 16-vCPU system behave similarly if the density is the same. This gives architects a portable way to compare capacity across instance sizes and processor generations.

The right density also depends on the business goal. Interactive copilots and user-facing assistants should favor lower density because response time matters. Batch workloads such as IT workflows can often run at higher density. This gives teams a practical way to tune fleets around service-level objectives and total cost of ownership.

Agentic AI requires a new form of observability: Average compute (CPU) utilization is a weak primary performance monitoring signal for agentic workloads. Agents often alternate between waiting for model responses and then doing short bursts of compute-intensive work. Because of that “bursty” pattern, average utilization can look acceptable even when those bursts are creating queues and slowing down the user experience. Task latency (P95) is a better leading metric. It shows when workflows are starting to wait, even before average task duration meaningfully degrades. A practical operating model is to alert on P95 latency first, then confirm the issue by looking at sustained task duration.

Scale out by default: Scaling out adds more systems, increasing total agent capacity, while scaling up adds cores or memory to a single system for agents with heavier compute bursts.

Our testing data showed that scale-out is usually the better default. That aligns with the fact that agents are typically semi-independent and have modest per-agent bursts, it improves overall performance, supports high availability, often lowers cost, and makes it easier to preserve the target agents-per-vCPU ratio as the platform grows.

Scale up when agents require heavier parallel compute, shared state limits partitioning, memory locality matters, or licensing constraints apply.

Consider business implications: Where will agentic AI create business value first? The organizations getting production-grade results are wrapping an automation layer around workflows that already have codified rules and measurable service levels: code creation, regression test farms, ticket triaging, market analysis, and security review.

The ideal enterprise persona for agentic AI is therefore not the experimental user chasing novelty; it is the accountable leader who must improve cycle time and productivity, protect service quality, enforce policy, and scale adoption with cost in mind. 

Agentic AI’s value comes from helping businesses complete real work across teams, systems, data, and processes. For enterprises, the priority is not just better model performance; it is creating a reliable environment where AI agents can support business workflows, improve productivity, operate within governance requirements, and scale as adoption grows.

In practice, success with agentic AI depends on the right foundation to deliver consistent outcomes, manage cost, maintain control, and move confidently from pilots to production.

This content was produced by Intel. It was not written by MIT Technology Review’s editorial staff.



❌