❌

Normal view

Coinbase runs 1,200 agents and just slashed its AI bill in half

Close-up of a server rack with rows of network cables connected to switches, illuminated by green LED lighting in a dimly lit data center.

Vercel CEO Guillermo Rauch and Coinbase CEO Brian Armstrong run very different companies, but they’re making the same architectural bet. Instead of building around a single AI provider, both are designing production systems that can route work across multiple models.

Rauch and Armstrong aren’t making this decision in a vacuum. Frontier models have become much closer in capability for everyday engineering work, open-weight alternatives have improved dramatically, and the price gap keeps widening. That makes it much easier to justify routing work across several models instead of committing to one. 

Trillion tokens, zero loyalty

In an interview with TechCrunch, Rauch said that Vercel now routes more than a trillion tokens a day across millions of deployments, and that the company is actively moving away from one-lab partnerships. Rauch’s point highlights that the model has become just one interchangeable component in a larger inference pipeline.

That’s a significant position from the CEO of a company that serves as deployment infrastructure for a huge share of the frontend ecosystem. Rauch is calling single-lab partnerships obsolete.

Rauch is calling single-lab partnerships obsolete.

Cheaper defaults, smarter routing

Armstrong is making the same bet, and the financial results state his case. Coinbase cut its internal AI spend by nearly half while overall token usage continued to grow, without imposing usage caps on engineers.

Their playbook basically runs on three core levers.

First, it’s an internal LLM gateway. Coinbase deliberately defaults its engineers to lower-cost open-weight models, specifically Z.ai’s GLM 5.2 and Moonshot AI’s Kimi 2.7. Engineers can still pull down a stronger model if a specific job absolutely demands it, but the pricing gap makes the default choice obvious. GLM 5.2 costs roughly $1.40 per million input tokens and $4.40 per million output tokens.

Compare that to Anthropic’s Opus 4.8, which sits around $5 for input and $25 for output. You are looking at a three- to six-times cost reduction per token. And it holds its own on major coding benchmarks, scoring 62.1 on SWE-bench Pro, compared to GPT-5.5’s 58.6. Plus, because Coinbase self-hosts these models, zero code or query data ever leaves their environment.

The second lever is task-based routing. Armstrong makes a highly practical point here, suggesting teams want a frontier model to do the heavy lifting for complex planning, but for pure execution tasks, where cheaper models perform just as well, there is zero reason to pay top dollar.

The third piece is aggressive caching. By keeping a conversation locked to the same model as long as the cached context is valid, Coinbase managed to push its cache hit rate from a measly 5% up to 60%. That 12x jump is a massive cost driver.

Gateways as control planes

If you want to understand Armstrong’s broader mindset, listen to his recent chat on the Sourcery podcast. He casually mentioned that Coinbase now operates with roughly 1,200 full-time AI agents, a number they calculate by normalizing compute hours to a standard 40- to 60-hour workweek. At that scale, he argues that human developers have absolutely no business manually choosing which model to use. The infrastructure has to automate that decision entirely.

Human developers have absolutely no business manually choosing which model to use.

Because foundation models are becoming so easy to swap in and out, the engineering focus is shifting to the surrounding infrastructure. Like a centralized control plane, a gateway intercepts every prompt and makes a dynamic, split-second decision about whether a workload actually requires the expensive reasoning capabilities of a frontier model or a cheaper, faster alternative can handle it. The infrastructure makes that call based on the cache state, the complexity of the task, and real-time pricing.

Teams need visibility into latency, uptime, token consumption, and cost across all providers because using multiple model providers changes observability requirements. Without that data, it’s difficult to know whether routing decisions are actually improving performance or reducing costs.

Test before you trust

Evaluation becomes just as important. Lower-cost models need to be continuously tested against the workloads that matter to an organization before they are deployed to production traffic. Public benchmarks are a useful starting point, but are no substitute for measuring how a model performs on your own code, data, and workflows.

Trying to pick the single best AI provider is a losing game.

What’s striking is that Vercel and Coinbase arrived at remarkably similar architectures despite solving different problems. Both assume that today’s best model probably won’t stay on top for long. If that’s true, the competitive advantage shifts away from the model itself and toward the infrastructure that decides which one to use. 

The post Coinbase runs 1,200 agents and just slashed its AI bill in half appeared first on The New Stack.

Watch AWS engineers troubleshoot agentic AI with OpenTelemetry and OpenSearch

A minimalist blue vector illustration of a person walking toward a massive, glowing open book that serves as a gateway, symbolizing the "bible" of data systems being rewritten for the future of AI and cloud-native architecture.

Your organization constantly needs more information about system performance, usage, and data while in production — or better yet, before it heads to prod. The challenge of telemetry increases with the complexity of your stack and agentic sprawl. Because “it works in the testing environment” becomes moot in the face of non-deterministic agents.

After all, AI agents span multiple environments, and that leaves traditional log-metric-trace models insufficient to handle the volume of the agentic AI era. The situation can lead companies to think that the best option is to throw everything into the locked box of proprietary tooling, but that creates another problem: Information is siloed within each layer, fragmenting data and taking you further from realizing real AI ROI.

Unified context across fragmented workflows

The OpenTelemetry framework and the OpenSearch distributed search and analytics engine make for a powerful, open-source pairing that gives organizations of all sizes unified context across their fragmented workflows. In fact, OTel has crossed the 95% adoption threshold for new cloud-native instrumentation projects and has already become the default choice for Greenfield projects.

OpenSearch, sponsored by Amazon Web Services, is gaining traction with AI engineers, as it recognizes that observability and AI must be united. This year’s OpenSearch roadmap specifically focuses on making it the primary retrieval interface for AI agents and an essential piece of any retrieval-augmented generation and agentic AI stack. 

Join us on July 22

Just because open source doesn’t have a direct cost doesn’t mean it’s free. That’s why Dotan Horovits and Rekha Thottan of AWS are going to perform a live troubleshooting simulation using correlated logs, metrics, and traces, followed by a demo of how agentic traces flow through Otel pipelines. Also learn how the open-source evaluation framework Agent Health can provide a structured pre-production benchmark to flag unpredictable agentic behavior before release. 

Join us live on July 22 to learn along and ask questions to learn how your organization can adopt these open-source standards in the second half of this year — across agentic workloads and traditional infrastructure, at scale.

Register for the webinar here

REGISTER NOW FOR THIS WEBINAR

The post Watch AWS engineers troubleshoot agentic AI with OpenTelemetry and OpenSearch appeared first on The New Stack.

The real cost, security, and culture problems behind enterprise AI agents

7 July 2026 at 20:24

Presented by Red Hat


At VentureBeat's recent AI Impact event, where the discussion centered on what separates enterprises that scale agentic AI from those that stall in pilot mode, Brian Gracely, senior director of portfolio strategy at Red Hat, detailed what companies actually run into once agents reach production.

He dove into cost discipline, the security blind spots unique to autonomous systems, and the organizational friction that determines whether agent adoption spreads beyond early champions.

Enterprises are overestimating how far behind they are on AI agents

Many enterprise leaders, especially those following industry keynotes and AI announcements, worry that they’re already falling dangerously behind competitors deploying agents at scale. But according to Gracely, much of that anxiety reflects a misconception about how quickly organizations learn once they begin building. Teams often move up the learning curve far faster than they expect.

That rapid progress creates a different challenge, however. As agent usage expands, AI costs rise just as quickly, turning cost management from an engineering concern into a recurring boardroom discussion.

Agentic AI usage is orders of magnitude higher than during the chatbot era, making AI costs a growing concern for enterprises. At the same time, organizations are becoming increasingly aware of their dependence on a small number of model providers. According to Gracely, that combination is driving many enterprises to explore alternatives that give them greater control over costs and infrastructure.

"The two or three top providers are already telling the market that they're losing money, and they're trying to go public to make up those gaps," he explained. "At some point, the dependency on that means you're either going to buy at a very high-cost level, or you're going to figure out alternatives to control what you're doing."

Right-sizing AI models is the fastest lever for cutting agent costs

The biggest cost issue is that enterprises overspend by defaulting to the most capable model available regardless of task complexity.

"If I'm simply trying to resolve an insurance claim, I don't need to know about the history of Western civilization in my model, I don't need to know World Cup soccer scores," Gracely said.

Semantic routing is the mechanism many companies use to make that judgment automatically, classifying requests and sending each to a model sized for the task without requiring users to choose, while infrastructure techniques like caching repetitive queries cut how often a request needs to reach GPU compute at all. Together, he said, these tools remove the assumption that efficiency and innovation pull in opposite directions.

"There's a lot you can do at a GPU infrastructure level, and quite a bit you can do in terms of flexibility of models," he explained. "Those give excellent choices in terms of the levers you're trying to pull, whether you need efficiency or you need innovation. That shouldn't be a binary choice."

The financial discipline needed for token spend is similar to the FinOps practices that took years to mature in order to take control of cloud compute spending. Those underlying frameworks will transfer even as the vocabulary changes, Gracely said, especially as organizations push for internal education on model selection so teams stop defaulting to the most prominent option for tasks that don't need it.

"The same way we first had to teach the financial people what an EC2 instance is and what an S3 bucket is, you're going to have to start explaining tokens to them," he said. "We don't always need a Rolls-Royce. We don't always need caviar, because we're trying to do basic types of things."

Patch speed is now critical as AI tools find vulnerabilities faster

AI-powered vulnerability discovery is forcing enterprises to rethink how quickly they can identify, validate and deploy patches. Long-established patch management cycles may no longer be fast enough in an environment where AI can uncover — and attackers can exploit — new vulnerabilities much more quickly.

"Most companies are probably going to have a window of somewhere between seven and 14 days to stay ahead," he said. "There are groups, Red Hat included, that are going to build patches for these, but the embargo window is going to be short."

AI is also changing what defenders need to look for. Rather than simply uncovering isolated critical flaws, AI security tools can identify combinations of seemingly minor vulnerabilities that become dangerous only when chained together. As both software complexity and vulnerability discovery accelerate, Gracely argued that the ability to rapidly manage and update software is becoming a strategic capability rather than simply an operational one.

Subject matter experts and compliance teams decide whether agents scale

In the end, organizational adoption comes down to the need for deep, sustained involvement from the subject matter experts whose knowledge the agent is meant to encode, which makes earning their buy-in a prerequisite rather than an afterthought.

"You have to think about the incentives, what you do for people who participate in this work so they don't feel threatened that it's going to take away their job, and how you incentivize people in the long run to cooperate with that innovation," he said.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Develop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00T

7 July 2026 at 17:05
As more teams move from humanoid robot bring-up to task-specific skill development, the need for repeatable development workflows is growing. Building humanoids...

As more teams move from humanoid robot bring-up to task-specific skill development, the need for repeatable development workflows is growing. Building humanoids remains complex, and today’s development pipelines are still highly fragmented. As a result, developers spend significant time configuring robotics infrastructure before they can focus on building robot capabilities.

Source

Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron

7 July 2026 at 17:00
Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context,...

Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context, determines the correct procedure, checks whether a specialist signal confirms the failure mode, and writes up a recommendation. This process remains consistent, and is well-suited for an AI agent. This post discusses a per-alarm…

Source

Maximize Spectral Efficiency with AI-Native RAN and NVIDIA AI Aerial

7 July 2026 at 17:00
An image of a 6G network.Spectrum is one of the most valuable assets in wireless communications. Over the last 30 years, telecom operators in the US have spent more than $240B to...An image of a 6G network.

Spectrum is one of the most valuable assets in wireless communications. Over the last 30 years, telecom operators in the US have spent more than $240B to acquire wireless spectrum. A goal of a radio access network (RAN) system is to extract the maximum spectral efficiency (bits/second/Hertz) possible, which translates into more capacity, stronger network resilience with fewer dropped packets…

Source

The AI architecture that let Liberty Mutual shrug off the Fable 5 outage

When Anthropic's Fable 5 was pulled from international use for nearly three weeks, some over-reliant businesses were left scrambling.

But Liberty Mutual easily pivoted to other platforms. That’s because 18 months earlier, they built their "AI backbone" exactly for this kind of scenario.

In this rapidly moving AI landscape, the 114-year-old property and casualty insurance company recognized independence as an operating advantage.

“Things are changing so fast, you need a backbone that's flexible,” Brian Craig, Liberty Mutual’s senior director of architecture, said at a recent VB Impact event. “You can't lock in right now on one vendor or even one framework.”

Enterprises need flexibility to hook into different models and vendors, depending not so much on the "flavor of the day," but “what can you feel confident about for the next six months,” he said.

Runtime versus control plane

The company’s "backbone" (or control plane) is its own, while everything underneath remains swappable.

The architecture consists of roughly 50 components across security, identity, orchestration, tool restriction, and the policies that govern how agents behave. Each is designed to be independently and immediately replaceable to support interoperability.

The agent runtime below this backbone is AWS's Amazon Bedrock AgentCore; this is not the strategic center, but explicitly “just for running the agents,” Craig said. He chose this offering because it (at least currently) supports multiple frameworks and Liberty Mutual’s model-agnostic philosophy.

“We still have flexibility based on what we write,” Craig said. “But if something comes along and is better, we will move to it quite quickly.”

The software factory

This architecture delivers, as proven by Liberty’s “software factory,” an agentic pipeline that automates much of the software delivery process.

They started with a business process with clear pain: Onboarding electronic content management documents for insurance products. This repetitive, manual task typically required engineers to code every change.

Instead, the team built a factory of coordinated agents working “in tandem and in sequence”:

  • An Epic agent consumes high-level requirements.

  • A Story agent breaks work dictated by the Epic agent into narrow slices within specific application areas. This agent is “constraining the context, because the smaller the context, the better the output.”

  • A planning agent defines the technical execution plan.

  • A coding/testing agent handles coding, testing, and basic review.

  • A triage (critic) agent sits across all other agents, reviewing quality and feeding back improvements.

  • Finally, a librarian agent helps others find the “context of the knowledge for their job.”

Craig and his team learned quickly that a single “do everything” agent was a mistake. “You were asking it to do too many things, which meant you had to give it too much information,” Craig said. Splitting into six agents let them dramatically shrink context windows and tighten scope.

Once the factory hit production, the impact was immediate. In the initial deployment, they did “about three months of work” in roughly a week. They realized that “the current software engineering process has a massive amount of handoffs, which means there's a massive amount of wait time,” Craig said.

Human-paced automation

The factory is not a fully autonomous pipeline; it runs at the speed of human overseers.

Liberty’s first model was a “day shift/night shift” rhythm: Engineers set goals and rules and reviewed the previous night’s outputs during the day, then let the factory run overnight. But in practice, there was never enough work to keep the agents busy all night, and the cadence felt unnatural.

They shifted to a more iterative loop. Users decide when to trigger the factory, how far it runs before pausing, and at which points they want to review outputs. “Then the factory kicked in, and it may only run for less than an hour, and then you would look at it again,” Craig said. “It was more controlled at the speed that our users felt comfortable with.”

The orchestration layer lets teams choose whether to review after the Epic stage, after planning, or once coding and testing complete. “That is up to the users of the factory, but it has completely removed a lot of the wait time that we currently have within our processes,” Craig said.

By contrast, early on, “every time the thing ran, they wanted to look at the output,” and that became the feedback loop that trained both humans and agents.

Some of this is, as Craig put it, “just automating agile at speed” and giving iterative feedback; some from humans, some from the triage agent. “But we're seeing that start to bake in, and it starts to become the rules.”

When the same feedback is coming from two different directions (agent and human), it makes sense to go into Liberty’s context repository. The agents will then rely on that context the next time they need to make that decision, "and the next time again, and you start to speed up."

“It's like a flywheel once you start building these and you start to get them flowing,” Craig said. “You realize it listens to what you say.”

Contracts that match the pace of change

Liberty paired that architectural posture with a contract posture, deliberately shifting from five-year enterprise deals to one-year agreements. The logic is simple as Craig sees it: The AI market moves too fast to lock into one vendor or framework for half a decade or more. Shorter terms let them evaluate — and if necessary, swap — models and platforms at the speed the market actually changes. Cost is part of the story. When a premium frontier model like Fable arrives, the sticker shock is real. “You see the price and go, ‘Goodness, it better be really good,’ Craig said. (It was; they got to use it enough to "fall in love with it.”) The backbone’s interoperability lets his team compare models at different price points and route workloads based on price–performance rather than vendor inertia. The same attitude governs how they plan to use agents from major SaaS platforms. Liberty is a customer of Salesforce and Splunk, who will both “bring their agents to the table.” His team has no interest in replicating engineering work, but they do insist on observability. “We just want to harness it as part of our system,” rather than control it, Craig said. “But we want the observability to understand what their agent is doing with our data, with our users.”

Closing the “control gap”

Importantly, Liberty built observability into its backbone. As Craig explained, it’s not just logging what an agent does, but what it accesses, which identity it uses, and which tools it’s allowed to invoke. Identity and access run in part on Microsoft Entra ID, and agents are given only the tools and permissions explicitly assigned to them. Whenever an agent realizes, ‘I don't have the information,’ it asks for it, and it only gets what it needs, rather than giving authority to “use every tool in the box.” “Because too much information given to an agent is worse than no information,” Craig said. “You just overload it, and it gets confused.” For detection, Liberty runs evaluations with MLflow against “golden datasets.” Whenever prompts or models change, they regression-test and immediately see whether results improved or degraded. One of his team’s new mantras is “you need to walk in the footsteps of a new start." If a new start can't find a guiding document, how will an AI agent? One of the key things enterprises need to do, no matter the business process, is “write stuff down, which is not earth-shattering,” Craig acknowledged. Agents obey written standards more reliably than people, and a context repository is a central artifact of Liberty Mutual’s system. Agents have made human judgment more, not less, central, he emphasized. Nothing ships without a human sign-off, consistent with Liberty’s risk posture as an insurer. “The confidence isn't there yet for us to just let it run wild, and I don't think it ever will be for the likes of Liberty,” said Craig. “We have to be rock solid before we let [anything] into production.” Moving fast for product fit and survival might make sense for other companies, but ultimately, “some people will get black eyes, but that's the joy of innovation these days,” he said.

Box survey: Why enterprise AI leaders are outperforming their peers

7 July 2026 at 16:25

Presented by Box


Content access, governance, and platform flexibility are emerging as the dividing lines between AI leaders and laggards, according to the new State of AI in the enterprise report from Box, which surveyed 1,640 IT decision makers across the US, UK, France, and Japan. One of the report's major findings is the speed of the shift: the combined share of organizations describing themselves as advanced or leading edge soared from 8% to 64% just over the past year, while the share calling themselves early stage or not yet started collapsed from 53% to just 9%. Eighty percent of organizations reported a notable return on their AI investment, defined in the survey as an improvement of at least 10%, and more than half saw measurable business impact within six months of getting a project approved.

The swing is largely due to how enterprises are now organizing their AI use rather than to any single technical breakthrough, says Olivia Nottebohm, COO of Box.

"We've moved from standalone experimentation that lived at the individual level into systematized, integrated agentic operations, agents that are in production and can be used in a repeatable manner," Nottebohm says. "That's where the impact is coming from."

Why AI leaders get higher ROI than early-stage companies

The divide between tiers is a matter of execution. Significantly, half of leading-edge companies reported AI-driven ROI above 25%, compared with just 11% of early-stage companies, with the advanced (33%) and developing (16%) tiers falling steadily in between. But Nottebohm says the real differentiator was not whether companies adopted AI, but how rigorously they integrated and managed it.

"What separates the leading edge is the operating muscle they've built: the right teams to deploy agents, formal governance to control them, and consistency in the content layer those agents work from," she explains. "Earlier stage companies are approaching it in a much more ad hoc, experimental way, letting people play around with it without the same intent or structured design."

Content access is the biggest barrier to enterprise AI ROI

Content, rather than model quality, is the defining bottleneck of 2026. Ninety-six percent of organizations say agents need access to company-specific content, yet only 36% have connected agents to trusted content across many use cases. It's an issue of trust rather than raw capability.

"We started this journey assuming enterprise AI was about access to the latest model," Nottebohm says. "But the question now is whether agents have access to the right content, and whether that content is protected, because those agents are only as good as the content they can reference, and only as safe as the security around it."

Getting that content layer right has a second benefit beyond safety, since it’s also what finally lets agents work across departments that previously operated in isolation from one another. And while roughly a quarter of organizations point to data fragmented across systems, 24% cite difficulty integrating AI into existing systems, 21% say they lack adequate permissions and access controls, and 18% describe their content as too unorganized to make accessible at all. Among the most mature organizations, 63% now treat unstructured documents, contracts, and reports as a competitive advantage rather than dead weight sitting in a digital filing cabinet.

Reducing common AI data exposure incidents

Nearly half of all organizations say they have already experienced an AI-related data exposure incident. That figure rises to 60% among leading-edge companies, which may face greater exposure from more agents and connected systems — but may also be better equipped to detect it.

The share of organizations reporting established or advanced governance frameworks rose from 24% in 2025 to 73% this year, but real gaps remain in instrumentation: only 39% have comprehensive visibility across sanctioned and unsanctioned AI use, 34% have formal standards for how agents access company data, and 27% still describe their governance as ad hoc. But those incidents function as a forcing mechanism rather than a setback, Nottebohm says.

"Governance used to be seen as something that slowed people down, but 93% of respondents told us better governance is actually what let them move faster," she explains. "It makes scaling AI survivable. Once content is secured and highly permissioned, you can run multiple agents across multiple processes and get a real multiplier effect."

One practical consequence of that shift is that permission structures built for human employees are now being revisited with agents in mind, a process most enterprises are only partway through.

"The permissions enterprises set up two years ago need to be reviewed," she explains. "Until fairly recently, people weren't setting permissions on a document with how an agent might use it in mind, but now they're much more deliberate about that. It leaves them with a whole corpus of unstructured data to go back through and either clean up or repermission."

That's part of a broader move away from governance designed for people and toward governance designed for agents from the start.

"Enterprises need to make the transition from governance that's retrofitted from human workflows to governance that's built specifically for agents," Nottebohm says. "That means tracking what an agent has touched, whose permissions were applied, and which sources were used, and all of that is now shaping how governance gets applied."

Enterprises need to avoid lock-in to a single AI vendor

"The days of token-maxing are already gone," Nottebohm says. "It's now about the responsibility of delivering efficient AI. Organizations want to use the cheapest model that meets the quality bar they need, not necessarily the most expensive one, because different model families keep leapfrogging each other and companies want to preserve that choice."

That means enterprises are avoiding lock-in more than ever. Sixty-eight percent say they're concerned about depending on a single AI provider, the average number of officially adopted AI tools has climbed to 3.3, and 79% now consider it important or critical that agents operate headlessly, connecting directly to systems and APIs without a human interface in between.

It's a trend similar to the shift toward multi-cloud infrastructure, and driven by a similar reluctance to hand any one vendor outsized negotiating power.

"A flexible architecture is built on platform interoperability," Nottebohm says. "It runs on multiple models, operates headlessly, and keeps every part of the AI stack swappable, so organizations don't have to bet on which individual tool wins, and that's part of the broader shift away from defaulting to the biggest, most expensive model available."

The next steps to AI success

Over the next three years, businesses should prioritize organizing, classifying, and cleaning up unstructured content, actively hiring and building teams around emerging roles, and adopting a hybrid token compute budget model, where IT owns the core infrastructure and token budget while business units own the application-level spend. And right now, it's easy to get up to speed fast.

"You don't have to start at early maturity and slowly work your way up," Nottebohm says. "If you build in the governance, the content layer, and the multi-model system from the start, you can enter as a leading company and capture that same outsized impact."


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Intelligence is Free, Now What? <br> Data Systems for, of, and by Agents

... government of the people, by the people, for the people ...
    — Abraham Lincoln, Gettysburg Address (1863)

The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; today the same runs under $1, and some providers are pushing costs below $0.10. Across benchmarks, inference prices have fallen between 9x and 900x per year, with a median decline near 50x. Even frontier models are getting dramatically cheaper each generation, with open-source models following closely behind. And crucially, even if “Nobel-Prize-winning genius-level” intelligence isn’t here yet, the intelligence that suffices for the vast majority of knowledge work is here today, and getting cheaper by the month. At this rate, we are soon entering the era of virtually free intelligence—the kind that is more than enough for everyday knowledge work.

A cartoon database character and an AI robot agent holding hands

Disclosure: This post is a perspective led by Aditya G. Parameswaran—an Associate Professor of EECS and co-director of the EPIC Data Lab at UC Berkeley—together with his collaborators. It is part landscape survey and part perspective, and several of the research directions discussed below (including agentic speculation, structured memory, and synthesizing custom data systems from scratch) draw on the authors' own ongoing work.

So, what does this new era of near-free intelligence mean for data systems? We believe three new challenges—and opportunities—stem from near-zero inference costs:

Data Systems For Agents. Agents will soon become the dominant workload for data systems—with swarms of agents spun up in response to each end-user request. Given differences in characteristics between agents and humans—or applications acting on their behalf—how should we redesign data systems for such agentic users?

Data Systems Of Agents. As agents start taking on the bulk of knowledge work, a new substrate is needed for thousands of agents to manage state over long-running tasks, coordinate and reach consensus, and deal with failures. What do data systems that reliably and efficiently run and manage agent swarms look like?

Data Systems By Agents. Agents are rapidly becoming capable of synthesizing entire data systems in one go—meaning we can rebuild custom systems for each new workload. Verifying that such systems match intended behavior is a challenge. What does it take to let agents synthesize data systems we can actually trust?

A database character and a robot agent holding up a triangle labeled 'of', 'for', and 'by'
Data Systems For, Of, and By Agents

Next, we will discuss each in more detail, followed by discussing the intertwined future of data systems and agents, especially as the three challenges intersect.

Data Systems For Agents

An agent querying a database doesn’t behave like a person or a BI tool. It performs what we call agentic speculation: a high-volume, heterogeneous stream of work spanning schema introspection, columnar exploration, partial and then full query formulation. With multiple agents each exploring portions of the hypothesis space, each user request could amount to 1000s of individual SQL queries. Now, users can issue ‘high-level’ data tasks, e.g., root-cause analysis—e.g., ‘why did coffee sales in Berkeley drop this year’—or exploratory cohort analysis—e.g., ‘which user segments are most likely to churn next quarter’—each involving a combinatorial space of potential joins, aggregations, and filter combinations.

An agent sending many SELECT SQL queries to a database and receiving results back
Data Systems Redesigned to More Effectively Support Agentic Speculation

The requests from these agents have various opportunities for optimization. For instance, on a text-to-SQL benchmark with multiple agents attempting each task, only 10-20% of the sub-plans are distinct. Thus, 80-90% of sub-queries perform duplicate work. The same experiments show task success rates significantly increasing with more agentic attempts—so the redundancy is actually helpful. But from the data system perspective it’s wasted work.

An agent-first data system can exploit such properties to help agents make progress faster. It can reuse results across overlapping sub-plans, drawing on ideas from decades-old literature on multi-query optimization and shared scans. Or the data system can try to satisfice, returning approximate answers that are good enough for agents to make progress, leveraging work from the AQP literature—or streaming the results of the final or intermediate operators to help agents decide if seeing the rest is necessary or helpful.

Another opportunity here is to rethink the query interface entirely: instead of agents issuing a single SQL query at a time, they could instead issue a batch of queries, each with its own approximation requirements. Since enumerating an exponential search space (as in the root cause or cohort analysis examples above) isn’t a good use of agentic reasoning ability, perhaps data systems should support higher-level primitives rather than requiring agents to list each SQL query explicitly. One idea here is to draw on DBT-style Jinja macros to provide looping-based primitives for agents to interact with data systems.

A swarm of AI agents working at laptops
A Caffeinated Army of Agents Ready to Tirelessly Complete Your Data Tasks

A final opportunity here is to stop thinking of data systems as passive executors of queries; data systems could be proactive, as they possess more grounding in data and system characteristics that agents may lack a priori—they could steer agents in different directions, provide results for related queries, and also provide performance-level feedback (e.g., instead of executing an expensive query, the system could first provide the agent a latency estimate). The reason we can do this now as opposed to the past is that an agent can accept any form of textual feedback and isn’t expecting a strict SQL query result. In fact, the data system could also prepare both materialized and virtual views for an agent in advance, provided to the agent as part of context, as this may be cheaper or more effective than having an agent author or use them.

Data Systems Of Agents

Previously, we focused on how agents interact with data systems. Now, we consider everything else agents need to keep working: where they live, how they remember, how they coordinate with each other, and how they deal with failures of each other. This agentic substrate is separate from the inference stack powering raw intelligence. However, the inference stack itself is being abstracted away through APIs (e.g., from OpenAI or Anthropic), or, for open-weight models, through serving frameworks that hide low-level details. So far, the agentic substrate has been managed through harnesses like Claude Code and Codex, coupled with various mechanisms to store and retrieve memory.

First, on the memory front, the current wisdom is that files are all you need; agents write to unstructured markdown (MD) files, which can then be searched using grep, or via embedding-based retrieval. In fact, many argue that the solution to continual learning is having agents consume a lot (e.g., an entire codebase, slack, company wikis, …) and then write their learnings into MD files, which are then retrieved selectively on demand. Indeed, file systems, bash scripting, and MD files are and will still be important for agents. However, at scale, when agents are doing the vast majority of knowledge work, this approach will no longer be effective.

Given limited context windows, retrieving all MD file fragments that may be relevant and stuffing it into the context will break down at some point. Even if context windows continue to grow, there are latency benefits to not put all information into context — and in many cases, e.g., when knowledge work involves interacting with large databases or code bases, it will be infeasible to serialize all relevant data into context.

A swarm of robot agents holding hands, each drawing state from a single large shared database platform below them
Data Systems As A Substrate for Multi-Agent Swarms

One could use a knowledge graph representation, but knowledge graphs suffer from the same limitations as unstructured MD-based memory due to their lack of structured search. What one needs is to be able to retrieve only memory that is pertinent to the task, across multiple attributes (or facets) of interest. For example, an agent debugging a flaky test should be able to pull only the memories tagged with the relevant module, language, framework, and failure mode—rather retrieving based on keywords or embedding similarity. A separate issue is what to actually retrieve; raw agent traces with mistakes are not very useful as they will induce agents to repeat the same mistake—instead, we want the retrieved memory to be corrective.

We recently explored a related notion of structured memory, where we organize memory across various attributes, each of which could be set as * to indicate universal applicability, or set as a list of values to be matched. For a data agent, the dimensions could include the columns and tables, type of operation, and finally, open-ended natural-language corrective instructions. So, we could include memory that only applies to a given type of operation (e.g., ‘when performing date-time operations, use fiscal year as opposed to calendar year conventions’), or a given table (e.g., ‘column product_cleaned is preferred over column product when querying on product name’). One open question is defining an application-specific structured memory—or what others have called world models for memory. We believe this is akin to defining a schema for each application—and perhaps agents themselves can help us define and refine it over time.

Diagram showing corrective knowledge stored with structured attributes (SQL keywords, tables, columns, data type) and retrieved by matching the features of a new agent query
One Possible Way To Store and Retrieve Structured Knowledge [From Here]

Structured memory will be useful also for evolutionary frameworks to effectively manage search spaces. Indeed, storing, structuring, and mining large volumes of single and multi-agent traces can help future agents become much more efficient—potentially enabling effective recursive self-improvement through structured memory-based mechanisms.

Another challenge is to support concurrent edits to shared memory, and concurrent edits in general, when there are many agents performing transformations. While there have been some useful attempts at supporting multiversioning and copy-on-write semantics, it isn’t clear that such techniques will suffice when thousands of agents are attempting to edit shared state at the same time. For instance, when agents are trying various potential transactions in response to a user request, the effects of the vast majority of these transactions need to be rolled back—with only the one ‘correct’ transaction’s result persisting. Work on supporting exactly-once semantics is relevant here, as are underlying techniques based on CRDTs and operational transformation. For updates to fuzzy mechanisms such as memory, we may be able to sacrifice on consistency for perfect correctness in the interest of latency. While agents can reason about semantics to compensate or roll back their actions to eventually finalize most tasks, the primary challenge lies in the degree to which they step on each other’s toes during the process. An important failure mode to be avoided is a form of “livelock,” where incessant compensating actions prevent any meaningful progress.

Beyond shared state, other concerns emerge when trying to support an army of agents, including what to do when agents fail, how agents should communicate with each other (directly or through intermediate shared state), and how we should deal with straggler agents. There have been some developments in supporting durable multi-agent execution, such as Temporal, but it remains to be seen if such solutions will apply at scale across thousands of agents. On the topic of communication, we need mechanisms to enable agents to negotiate with each other. Imagine four developer agents attempting to reach consensus on a shared schema, with distinct but overlapping objectives. In a human setting, this would involve iterative discussion and compromise; for agentic swarms, we must define the mechanisms that allow them to converge on a design that reflects the underlying goals of their respective principals. Or if agents are all requiring access to a limited resource, again communication will be necessary. It remains to be seen if this is best done via centralized coordination, or if a decentralized approach is necessary.

Data Systems By Agents

Finally, if intelligence is effectively free, then we can employ this intelligence to synthesize new data systems from scratch. Indeed, in many settings, general-purpose data systems may be overkill, as they have to support every schema, query, and hardware target. Given a workload, recent work, including Bespoke OLAP and GenDB, has shown that one can use an agentic pipeline to synthesize a complete, workload-specific analytical engine—in minutes to a few hours, at a cost of a few dollars. The engines are disposable: when the workload shifts, one can simply regenerate them. Analogously, our work has shown that one can synthesize custom key-value stores from scratch, targeted to the workload. In fact, modern IDEs, such as Kiro, elevate specifications for systems development to be a first-class citizen.

A robot agent with a hammer and chisel carving a database character out of a block of stone
Agents Can Synthesize Custom Data Systems From Scratch

The main issue, however, is that specifications are typically imperfect, and don’t cover all corner cases. Present-day agents will exploit the missing specifications to reward-hack their way to a high performance metric. In our custom key-value store work, we found that one way to alleviate this is to have auxiliary verification agents trying to generate test cases that catch the exploitation of corner cases, essentially expanding the specification. Yet another approach is to both generate a system and a proof for its correctness together, for which we have found some early success, but more needs to be done to solidify the approach. Further, it remains to be seen what is the best way to solicit human-written specifications for a system—can this be done in an iterative, human-in-the-loop manner, as opposed to a one-shot, incomplete one. Indeed, human-written specifications are incomplete even for manually authored software, so one would expect that future agents that are more aligned will increasingly exercise better judgement when making design decisions.

Pipeline diagram where a system builder provides a specification, planner and coder agents generate code, the code is evaluated for correctness and performance, and critic and auditor agents provide feedback and catch reward hacking
One Possible Data System Synthesis Pipeline [From Here]

Other questions here involve testing whether starting from a mature system (e.g., Postgres) and removing components/functionality can lead to higher performance or more user trust. Separately, is there an opportunity to make the design composable, comprising various verified components that are mixed and matched given a workload? For example, perhaps the workload hasn’t changed enough for the storage layer to be updated, but perhaps the query optimizer requires changes. A perhaps more viable proposition involves employing agents coupled with proof systems to target critical parts of the code associated with formal proofs, rather than doing so for the entire system.

A final opportunity here is to move away from the traditional data systems stack with clearly-defined interfaces (e.g., parser, query optimizer, storage manager, …) — that were each largely the prerogative of a single human team to manage. Instead, agents can find new ways to “blend” these components together, perhaps identifying new optimization opportunities as a result. Agents can also fill in missing gaps in functionality to make existing systems much more feature-complete, or reach feature-parity with other competing systems—or analogously, continuously refining open-source systems in response to feature requests or issues (perhaps filed by other agents!) Doing so in a way that prioritizes correctness, long-term maintenance, and human interpretability will be a challenge.

Looking Further Ahead

In the era of near-free intelligence, data systems matter more than ever. As agents take on the bulk of knowledge work, the workload for data systems will change, the substrate they need to run on will have to be built, and increasingly, they will participate in designing data systems themselves. Each of these shifts opens up a new, exciting research agenda.

A half-database, half-robot character next to a yin-yang symbol formed by a database and a robot agent
Co-Evolution of Data Systems and Agents

Looking further out, the boundaries between agents and data systems will likely start to blur. For instance, agents may design the data systems they themselves run on, defining both the interfaces as well as the system components underneath. Both the interfaces and internals can be evolved over time by agents in a form of recursive self-improvement. There is also an opportunity to rethink data systems as a holistic source of truth for the entirety of relevant state: including raw data, memory, and coordination state, further erasing the distinctions between the data that is being queried by agents and data generated as a result of agentic activity. Finally, data systems may themselves incorporate agentic components, fundamentally evolving from passive computation engines into intelligent, proactive, self-optimizing architectures. It is hard to predict what the future may hold. We’re in for a wild ride!

Acknowledgments

The perspective and ongoing work described in this post are the product of joint research and many discussions with wonderful collaborators at the EPIC Data Lab, Data Systems & Foundations group, and the broader Berkeley AI-Systems community. Thank you all!

BibTex for this post:

@misc{intelligence-is-free-blog,
  title={Intelligence is Free, Now What? Data Systems for, of, and by Agents},
  author={Aditya G. Parameswaran and Shubham Agarwal and Kerem Akillioglu and Shreya Shankar
          and Sepanta Zeighami and Rishabh Iyer and Matei Zaharia and Alvin Cheung
          and Natacha Crooks and Joseph Gonzalez and Joseph Hellerstein and Ion Stoica},
  howpublished={\url{https://bair.berkeley.edu/blog/2026/07/07/intelligence-is-free-now-what/}},
  year={2026}
}

Anthropic brings Claude Cowork to mobile and web as usage data shows most users aren’t coding

Anthropic on Tuesday launched Claude Cowork on mobile and web, expanding a tool that has quietly become the company's bridge between the developer-centric world of AI coding agents and the far larger market of knowledge workers who never open a terminal.

The rollout, which begins in beta with Max subscribers before expanding to additional plans, marks a strategic inflection for Anthropic. It transforms Cowork from a desktop-only agent into a cross-device platform where tasks can start on a laptop, continue autonomously in the background, and be reviewed from a phone — even after the user closes the app entirely.

"Your work goes everywhere with you, and keeps going without you," Anthropic writes in its announcement.

The timing is deliberate. Alongside the mobile launch, Anthropic published usage data from 1.2 million anonymized Claude Cowork sessions sampled between May 11 and May 31, drawn from more than 600,000 organizations. The data paints a striking picture: the overwhelming majority of what people do with Cowork has nothing to do with writing software.

The biggest AI story nobody's talking about

The numbers tell a story that cuts against the dominant narrative in enterprise AI, which has fixated on coding assistants and developer productivity as the primary use case for large language models.

Business process and operations — tasks like pulling scattered updates into a single report, building onboarding checklists, and reconciling spreadsheets — accounted for 33.4% of all sampled Cowork sessions, making it the single largest category by a wide margin. Content creation and copywriting — producing drafts, slide decks, posts, and proposals — came in second at 16.4%.

Together, those two categories make up roughly half of all Claude Cowork usage. Software development, by contrast, accounted for just 8.7%. DevOps and infrastructure followed at 7%, with research and intelligence at 6.4%, data analysis and business intelligence at 5.8%, document processing and extraction at 4.1%, and sales and revenue operations at 4%.

The remaining 12 categories each represented less than 4% of usage, including personal assistance at 3.8%, education at 2.4%, and meeting intelligence at 1.8%.

Anthropic describes these dominant use cases as "the work around the work" — tasks that span nearly every role in an organization but rarely appear in anyone's core job description. "People are using it for a variety of tasks that aren't necessarily the hallmark of a specific role, but instead represent the connective work around a role that moves projects forward and keeps businesses running," the company writes. "That means tasks like drafting a status update, building a slide deck, or condensing reams of research into a single report."

That phrase — "the work around the work" — is Anthropic's attempt to define and claim an entirely new category of AI productivity. It's a calculated reframing: rather than positioning AI as a tool that replaces what professionals do, Anthropic is arguing that the most valuable current application is handling everything professionals do around their actual expertise.

What mobile access changes — and what it doesn't

The expansion to mobile and web introduces three concrete capabilities that reflect how Anthropic envisions Cowork fitting into daily workflows.

First, sessions now sync across devices. A user can start a task at their desk, check on its progress from a phone, and retrieve the finished output from any device. Second — and arguably more significant — Cowork can now run tasks in the background with no device online at all. Users can schedule work for a specific time, and Claude will execute it autonomously. Anthropic offers the example of setting Monday morning client prep for 6 a.m.: "Claude works through the email threads, transcripts, and recent news, builds the briefing doc, and leaves the follow-up email drafted but unsent. Review it over coffee."

Third, when Claude encounters a decision that requires human judgment, it surfaces the question to the user's phone. "Nothing ships until you've reviewed and approved it," Anthropic states.

Desktop remains the most fully featured surface, with access to local files and the browser. But the web version also opens Cowork to users who cannot install a desktop application — a meaningful expansion in enterprise environments where IT departments control software installation.

The company also unified its interface: on web and desktop, chat and Cowork now share a single home screen, and projects and artifacts persist across both modes.

To encourage adoption, Anthropic is extending doubled Cowork usage limits through August 5.

The strategic logic: why Anthropic is chasing the non-developer

The usage data and the mobile launch together reveal a company executing a two-track strategy. Claude Code, its terminal-based coding agent, dominates among software developers. But Cowork is designed to capture the vastly larger population of professionals whose work involves creating, organizing, and communicating information rather than writing code.

The contrast between the two products is instructive. As Anthropic notes, Claude Code "is most often used by software developers for the key parts of their role: building, debugging, and shipping code." When developers do use Cowork, they tend to use it not for programming but for the communications-focused work that surrounds every role — status updates, documentation, and coordination.

This pattern — where AI handles the connective tissue of work rather than its core substance — aligns with what Anthropic describes as people using "Claude Cowork to assemble and structure the information they can use to act on their expertise." The company illustrates this with three examples: a lawyer using Cowork for document formatting and filing while reserving legal judgment for themselves, a hiring manager synthesizing interview feedback while spending more time on candidate conversations, and a team lead producing a slide deck that explains a decision while focusing on actually making that decision.

The implications for Anthropic's business model are significant. Developer-focused tools, while high-profile, serve a relatively narrow market. The Ramp AI Index published in May showed Anthropic pulling ahead of OpenAI in business adoption for the first time — with 34.4% of firms paying for Anthropic's services compared to OpenAI's 32.3% — and suggests the company's enterprise push is gaining traction. Claude Code was identified as the primary driver of that shift. But Cowork targets an addressable market that is orders of magnitude larger: every knowledge worker with a laptop, a pile of spreadsheets, and a slide deck due by Friday.

A crowded field gets more competitive

The mobile launch arrives during one of Anthropic's busiest — and most turbulent — stretches in its history.

Just last week, Anthropic launched Claude Sonnet 5, a new model that narrows the performance gap with its more expensive Opus-class models while maintaining lower pricing. The model is available at introductory pricing of $2 per million input tokens through August 31 before rising to $3 per million input tokens. Sonnet 5 serves as the engine underneath Cowork, and its improved agentic capabilities — better reasoning, tool use, and sustained task completion — directly enhance Cowork's ability to handle complex, multi-step workflows.

Two weeks before that, Anthropic released Claude Tag, a Slack-native AI agent designed for team collaboration. Where Cowork focuses on individual task delegation, Claude Tag operates as a multiplayer tool — a single Claude identity that everyone in a Slack channel can interact with, building context from conversations over time. 

According to Anthropic's announcement, 65% of the company's own product team's code is created by its internal version of Claude Tag. Fortune reported that Anthropic's head of product for Claude Code and Cowork, Cat Wu, described the distinction: "Claude Code, Cowork, and chat are very single-player, whereas Claude Tag is built to be interactive and multiplayer."

Together, Cowork and Claude Tag represent a pincer strategy: Cowork captures individual productivity workflows across devices, while Claude Tag embeds AI into team communication channels. Both are designed to push Anthropic deeper into enterprise operations, beyond the developer seat.

The security question looms

The expansion also arrives against a backdrop of unresolved security concerns. On July 1, security firm Armadin — led by Mandiant founder Kevin Mandia — published research detailing what it described as a full sandbox escape in Claude Cowork on Windows, as reported by SiliconANGLE. The attack chain involved DLL sideloading against the Claude desktop executable to gain trusted access to Cowork's virtual machine service, then exploiting undocumented parameters to achieve root access and bypass network restrictions.

Anthropic responded that the vulnerability did not qualify as a security issue because exploiting it requires an attacker to already have local code execution on the host machine. Armadin, however, raised a broader concern: that deploying local virtual machines on nontechnical users' systems creates visibility gaps that endpoint security products struggle to monitor.

This tension takes on new dimensions as Cowork moves to mobile and web. The web and mobile versions run tasks server-side rather than in a local virtual machine, which eliminates the specific attack surface Armadin identified but introduces different questions about data handling, especially for scheduled background tasks that process email threads, calendar data, and documents without real-time user oversight.

Anthropic's announcement states that "the decisions still come to you" and that nothing ships without review and approval. But as Cowork takes on increasingly complex autonomous workflows — processing contract folders, building client briefings from multiple data sources, drafting emails — the surface area for prompt injection and data exposure grows correspondingly. 

When Cowork first launched in January, TechCrunch reported that Anthropic explicitly warned about prompt injection risks, noting in its blog post: "These risks aren't new with Cowork, but it might be the first time you're using a more advanced tool that moves beyond a simple conversation."

As Anthropic courts enterprises, geopolitics complicates the pitch

Anthropic's enterprise push is also colliding with geopolitical reality. CNBC reported Monday that Alibaba will ban employees from using Anthropic's AI tools starting July 10, placing Claude Code on a high-risk software list. The move followed Anthropic's June letter to the U.S. Senate accusing Alibaba of carrying out what it called "the largest known distillation attack" against its models.

The Alibaba ban, combined with reports that Anthropic is closing loopholes that allowed Chinese companies to access Claude through third-country entities, underscores the increasingly fraught environment for AI companies attempting to serve global enterprise customers while navigating U.S. export and security restrictions.

At the same time, Anthropic is investing massively in infrastructure. Reuters reported Monday that Anthropic signed a $19 billion, 20-year lease with TeraWulf for a data center being built in Hawesville, Kentucky, with 401 megawatts of computing power expected to become fully operational in 2028.

That kind of capital commitment only makes sense if the company expects enterprise demand — not just from developers, but from the millions of knowledge workers that Cowork targets — to grow dramatically.

Anthropic's own usage report comes with notable blind spots

Anthropic is transparent about the limitations of its usage analysis. The taxonomy classifies sessions by the type of work being performed, not by the job title of the person doing it. 

There are no standalone categories for marketing, finance, or HR — functions that are likely absorbed into the dominant "business process and operations" bucket, which may partly explain why that category commands a third of all usage.

The sample is also rate-capped rather than proportional to traffic, meaning the numbers are shares of sampled sessions, not absolute volumes. Usage during peak hours is somewhat underrepresented. And roughly 5% of sampled sessions involved personal, non-work use — hobbies, personal assistance, and companionship-style conversations — meaning the data doesn't purely reflect workplace activity.

The company also acknowledged that its labeling pipeline changed around May 11, which is why the analysis window begins on that date rather than covering a longer period.

What Cowork's rise says about the future of enterprise AI

Anthropic's mobile launch and usage data arrive at a moment when the enterprise AI market is shifting from proof of concept to proof of value. The question facing every company deploying AI tools is no longer whether the technology works — but whether it delivers measurable productivity gains across an organization, not just within engineering teams.

The usage data suggests that the answer, at least for Cowork, is emerging in an unexpected place. It's not in the glamorous work of building software or conducting research. It's in the unglamorous, universal labor of turning messy information into structured outputs that move organizations forward — the status reports, the onboarding checklists, the variance memos, the client decks.

By untethering that capability from the desktop and making it available on every device, Anthropic is betting that the most valuable AI agent isn't the one that writes code. It's the one that handles everything else.

NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads

9 July 2026 at 18:10
Vera CPU image.Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration, and...Vera CPU image.

Agentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration, and result handling. As these systems scale across the AI factory, performance depends not only on GPU acceleration, but also on the CPU work that happens between model steps. Across the creation and deployment of an agentic system…

Source

Digital-native startups are ditching rigid databases for their agentic stacks     

7 July 2026 at 07:00

Presented by MongoDB


The gap between what AI models and agents can produce and what legacy infrastructure can reliably support is known as architectural drag, and it is the defining bottleneck of the agentic era. 

The data layer underneath an agentic system must handle variable schemas, vector embeddings, real-time retrieval, and multi-tenant scale, often simultaneously and without human intervention to manage migrations — but traditional relational databases weren't natively designed for document flexibility or AI capabilities. Fixed schemas require manual updates every time an AI agent introduces a new data shape, while separate vector databases add latency and synchronization overhead.

Three digital-native startups — Huntr, Modelence, and Tavily — solved this problem the same way: by building on MongoDB Atlas, a unified database platform with native vector search, hybrid search, and managed autoscaling. Their experiences define what an agent-native data stack looks like in production, and why using Atlas enables developers to easily build complex AI native companies.

Modelence: Building the agent-native cloud

Modelence is an AI app builder with an open-source framework designed specifically for agent-native development, enabling anyone to build and deploy production-ready web applications, including APIs and databases, in minutes. The company recognized early that most backend infrastructure was built for humans, not AI, and that the rigid schema management and complex migrations of traditional systems create operational drag that causes agents to fail when trying to build production-ready apps.

“Choosing MongoDB helped us keep everything in a single place, which is an important property of what we strive to do for our own users," says Aram Shatakhtsyan, co-founder and CEO of Modelence. "Live data streams, vector search, all as part of the main database. For AI agents, it’s especially important to have a single platform where everything can be done, because connecting multiple platforms together makes it more error prone.”

Modelence standardized on MongoDB Atlas because its document model aligns with how AI agents process and generate data, allowing schemas to evolve rapidly without manual migrations. The platform pairs that flexibility with a typed schema layer on top, a deliberate architectural decision. 

“MongoDB’s document model enables us to both keep things simple and at the same time decide how structured we want everything to be," Shatakhtsyan says. We still add a typed schema on top, which tremendously improves the accuracy at which AI can generate fully working, reliable web apps."

The TypeScript integration has been especially consequential, he adds. 

“Because MongoDB types and values can be directly translated to TypeScript, it becomes an extension of the Modelence framework and our App Builder has a single source of truth for both app logic and database,” Shatakhtsyan explains.

The result is a platform that can move from planning to a running live feature in minutes with significantly fewer regressions. That speed and reliability helped Modelence raise $3 million in seed funding and successfully launch an AI-native app builder that handles the entire application lifecycle end-to-end.

Tavily: The web access layer for agents     

Tavily is the search API purpose-built for AI agents, connecting them to real-time, accurate web knowledge and keeping them grounded in what's actually happening, not in static training data. At Tavily's scale, every agent request authenticates, retrieves, and meters without friction. That demanded backend infrastructure built to absorb change without breaking.

“On the user side, every agent request authenticates and meters against it," says Tomer Weiss, Data Team Lead at Tavily. "On the data side, we use it to track the lifecycle of every document we’ve ever touched: when it was fetched, how stale it is, what the freshness signals were and how popular it is. MongoDB’s flexible schema let us keep evolving those records without migrations as new metrics and features came along.”

That living record is what keeps agents grounded in reality. Multi-tenancy at Tavily's scale means managing millions of API keys, distinct usage profiles, plan tiers, and regional residency requirements. They built for that complexity from day one. 

“We separated concerns across clusters early: a user/account cluster optimized for low-latency authentication and usage writes, and a sharded cluster for document state where the scaling axis is URLs, not users," Weiss explains. "That separation has paid off.”

The most critical lesson is about choosing infrastructure that doesn’t punish change, and that flexibility compounds, he says. 

"The AI space moves so fast that change is our norm," he explains.  "For a company serving AI agents, where the workloads themselves keep changing shape, choosing a data platform that doesn’t punish change has turned out to be more valuable than any single feature.”

Huntr: From job tracker to AI career platform

Huntr.co, an AI resume building and tailoring platform, helps more than 500,000 job seekers across 190 countries craft stronger applications and manage their search. For a lean, three-person engineering team, the challenge was finding a data foundation flexible enough to store the full complexity of a person’s career history in a structure that AI could read, reason about, and generate from natively.

“The kinds of career data we are gathering at Huntr naturally aligns with MongoDB’s document model," says Trevor McCann, senior software engineer at Huntr. "The core problem we’re solving with AI job search tools is how to surface the qualities of a candidate that make them unique. We need to be ready to store whatever kinds of data the candidate wants to include in their materials.”

Huntr built its AI Resume Builder on MongoDB Atlas, where the document model mirrors the natural shape of career data: deeply nested, variable across candidates, and constantly evolving as the platform ships new features. MongoDB Search on Atlas handles core search needs while MongoDB Vector Search powers the Job Tailoring feature, which puts a candidate’s stored career profile side by side a specific job description and uses semantic matching to generate a resume optimized for that role.

The integrated capabilities have had a direct impact on how quickly the team can ship, McCann says. 

“MongoDB’s hybrid search allows us to seamlessly query across literal and semantic text matches, a must-have when working with such diverse data,” McCann says. “This is something we could piece together using other solutions but with MongoDB it’s ready to go on top of our existing data layer.” The consolidation of database, search, and vector capabilities into a single platform is what allows the team to punch above its weight. Huntr considers MongoDB the fourth member of its engineering team, McCann adds. 

Looking ahead, the platform is building toward AI that learns from a candidate’s full professional history over time, delivering more personalized guidance with every interaction.

The digital native blueprint

These success stories become a definitive "digital native blueprint" for the agentic era, built on three core pillars. First, by unifying database, search, and vector storage into a single platform, these startups have effectively eliminated the architectural tax of complex data schemas that typically slows down development. This consolidation enables a level of fluidity that is now non-negotiable; AI agents require a modern data platform that can adapt as quickly as a natural language prompt evolves. 

The winners of the AI era will be the ones who build the most performant, durable, and flexible systems to support those models in production. As agentic workflows grow more sophisticated, the data foundation determines how fast a team can ship, how reliably agents can operate, and how quickly the platform can adapt when the landscape shifts again. 


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Build for the new AI era with Microsoft and NVIDIA

6 July 2026 at 07:00

Presented by Microsoft and NVIDIA


Every generation of leaders has its own business transformation challenges to face. A decade ago, modernization meant cloud migration. Five years ago, it meant enabling remote and hybrid work. And just a few short years ago, the generative AI boom prompted organizations globally into enterprise AI adoption.

The demo era is ending

In the years since AI became the big new buzzword, generative models proved to be a crucial stepping stone, but the path to Frontier Transformation is agentic AI. Machine-generated answers aren’t enough when what the business really needs is sophisticated AI that can act. Experimenting with agentic capabilities was a necessary step; prototypes and pilots have proliferated. But the chapter on demos is closing.

Scale acceleration is beginning. In 2026, organizations that want to see agentic results that impact the bottom line must move from the knowledge layer to the action layer. And they certainly intend to—according to Deloitte’s 2026 AI report, 54% of enterprises surveyed expect to move 40% or more of their AI experiments into production. How hard could it be?

Agents are a different engineering problem—here’s why

Moving past prototype is the hardest part. Shipping an agent to production isn’t just a harder version of shipping a generative AI chatbot. Agentic production is a different engineering problem altogether, requiring orchestration, memory, runtime isolation, and ground-up observability—all required to deliver an agent that reasons, acts, and collaborates.

Once an agent moves into production, every tool and data source becomes an integration challenge. Running the agent requires isolation between sessions, durable state, and runtimes that hold up under a working load. And operational blindness turns agentic assets into liabilities. Once an agent is live, you need the ability to monitor, understand, and troubleshoot its systems across its lifecycle—a whole new discipline of observability is required, but teams don’t know how to get there. But we’ve been here before. When microservices faced a similar crossroads a decade ago, the lesson was this: those who recognized the need for a platform approach are the ones with the best success.

The production gap: Why most agent projects stall before scale

Moving from demos to real-world deployment introduces a host of challenges: how to chain multiple steps together reliably, how to ensure security and identity across agent components, how to monitor and improve agent behavior, and more. Many teams attempt to address these challenges with custom scaffolding, but the risk is often greater than the reward—slower time to value, gaps, and unreliability.

This is where the platform approach comes in. Without shared context and intrinsic trust, AI is difficult to rely on and hard to scale, with data fragmentation keeping production agents from matching pilot performance. Agents lack business context, enterprise signals are fragmented, development is complex and brittle, and security and governance are bolted on.

The solution is a unified platform that empowers developers to build, run, and scale agentic and physical AI end-to-end. Together, Microsoft and NVIDIA partner to enable this platform approach, helping enterprises effectively take agents from pilot to production.

What an agent factory actually looks like

Frontier Firms are those that not only successfully take agents into production but that also understand monolithic agents aren’t enough—a system of collaborative agents is key. They are the ones building agent factories, operating on a production philosophy that utilizes a reliable foundation and repeatable process for cross-functional, collaborative agentic solutions at enterprise scale.

So what is an agent factory? It’s a coordinated production architecture that combines an agentic control plane with accelerated specialist models, agents, and skills, allowing organizations to enable a governed system of models and agents at enterprise scale.

Within this production system, Frontier Firms are building heterogenous systems of agents, where the right models, tools, skills, and specialist agents are appropriately orchestrated at the right step of every job. The result is broad-reasoning frontier agents that plan, synthesize, and collaborate with users and other agents while accelerated specialist models and agents execute domain-specific work with speed and efficiency.

Microsoft and NVIDIA jointly empower this agentic factory approach. Microsoft delivers the enterprise control plane enabling runtime, identity, governance, observability, data access, and tool connectivity that agents need to collaborate safely. NVIDIA delivers the intelligence, acceleration, and specialist layers that give enterprises a repeatable way to move from isolated demos to governed, scalable agentic systems that can work together across business processes to accomplish meaningful tasks, not just answer questions.

At Microsoft Build 2026, Microsoft and NVIDIA showed how this architecture is coming together across cloud, local, and developer environments, bringing NVIDIA models, blueprints, and tooling into the Microsoft ecosystem to enable systems of agents with governance and speed:

  • NVIDIA models are now on the hosted agents in Foundry Agent Service.

  • NVIDIA’s open model portfolio on Foundry now spans agentic, physical, and scientific AI.

  • NVIDIA Agent Toolkit and NVIDIA NemoClaw blueprints give developers an open-source platform to build production agents on Foundry.

  • Foundry Local on Azure Local is now on the NVIDIA RTX PRO 6000 Blackwell Server Edition platform.

  • NVIDIA OpenShell integrates with GitHub Copilot for secure agent development.

You can read more about these announcements here.

Where to go from here

The organizations that win with agentic AI will be the ones that invest in a factory approach. Ready to take the next step on your agentic journey? Explore these resources:


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

The First AI‑Designed Vaccine Has Been Tested in People. Here’s What Happened.

7 July 2026 at 14:00

Scientists used AI to find targets shared by thousands of related viruses and build what they hope is a universal vaccine.

Researchers at the University of Cambridge have developed what they describe as a fundamentally new type of vaccine using artificial intelligence. The vaccine’s key component was designed entirely by AI and has now been tested in people for the first time.

The goal is ambitious: a single vaccine that works not just against all known human coronavirus variants, but against related bat viruses that could jump from animals to humans and cause future pandemics.

Traditional vaccines train our immune system to recognize one specific virus. The problem is that viruses mutate. When they change enough, the vaccine stops working, which is why we need a new flu shot every year and why Covid vaccines have been updated repeatedly since 2021.

AI offers a way around this. By analyzing genetic data from thousands of related viruses, it can identify the parts that stay the same across different strains and that are unlikely to change over time. Target those stable features, and you have a vaccine that should work against the whole family, not just the strain you started with.

This is exactly what the Cambridge team did. They used AI to scan viruses from the sarbecovirus family, which includes the viruses that cause both SARS and Covid, as well as a range of animal coronaviruses—looking for shared features that evolution has left largely untouched. Those features became the basis of the vaccine.

DNA Vaccines

While many people are familiar with the mRNA shots used during the pandemic, this new vaccine uses DNA. DNA vaccines are generally more stable than mRNA vaccines, making them easier to store and transport. This is a significant advantage in lower-income countries where “cold-chain” infrastructure is limited.

They can also be administered without needles. A high-pressure stream of liquid delivers the vaccine through the skin, making administration less painful and easier to scale up during an outbreak.

Could It Protect Against Future Pandemics?

These practical advantages matter most if the vaccine itself can do something no existing jab can: protect against viruses we haven’t encountered yet.

Broad-spectrum vaccines could change the way the world responds to emerging infectious diseases. By offering much wider protection than traditional vaccines, they could provide rapid immunity against new and emerging viral threats. This would equip public health officials with tools to stop future outbreaks in their tracks before they have a chance to turn into global pandemics.

They could also transform our approach to more familiar diseases. Influenza is a prime target because it exists in many different strains and evolves so rapidly. Scientists have to predict which strains will dominate each flu season, and if they guess wrong, vaccine effectiveness can suffer. A universal flu vaccine that targets features shared across multiple strains could eventually end the annual race to keep up with the virus.

The Ebola virus shows why this matters right now. The recent outbreak in the Democratic Republic of the Congo and Uganda is driven by the Bundibugyo strain, which bypasses existing vaccines. While researchers rush to create a new vaccine specifically for this strain, local communities remain at high risk. A broad-spectrum vaccine designed to cover an entire virus family could transform that picture.

What the Trial Found

This is the first human trial of an AI-designed vaccine. The results showed that this DNA vaccine was able to stimulate the immune system to produce antibodies that can recognize different types of sarbecoviruses. The technology was found to be safe and well tolerated.

This is an exciting advance because it demonstrates how AI has the potential to design variant-proof vaccines against future pandemic threats. The needle-free delivery system could also make the vaccine easier to administer and distribute worldwide.

However, there is more work to do. Although the results in this study are encouraging, the immune responses following vaccination were modest. It was also uncertain how long the protection lasts and whether further boosters will be required. Larger trials are also needed to determine whether the vaccine can prevent or reduce viral infections in the real world.

A universal vaccine remains a few years away. And any new vaccine must still pass larger trials to prove it is safe, effective, and provides lasting protection. But this study shows the goal is getting closer—and AI may help us get there faster.The Conversation

This article is republished from The Conversation under a Creative Commons license. Read the original article.

The post The First AI‑Designed Vaccine Has Been Tested in People. Here’s What Happened. appeared first on SingularityHub.

The foundational elements of AI architecture that IT leaders need to scale

With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces risk, leaving IT leaders to wonder which investments will prove valuable even six months into the future.

Returning to the foundational elements of AI architecture—the structural framework required for deploying and managing reliable, integrated AI systems at scale—allows technology leaders to make astute decisions today while supporting a future of AI agents that can retrieve information, make decisions, and execute complex workflows across systems.

Four elements of AI architecture you can count on

The following capabilities provide a stable compass on the path to production-ready deployment, regardless of how the underlying technology evolves.

1. Prepare data for AI at scale

Models are only as reliable as the data they can access, and poor data quality leads to AI hallucinations, bias, and unreliable outputs.

Most enterprises rely on legacy systems, inconsistent data structures, fragmented ownership, and incomplete datasets, making it difficult to scale AI effectively. Powerful as it is, AI itself cannot solve these underlying data problems.

As Adnan Adil, CIO of Elastic, explains: “The data is a durable part of AI architecture because without it, these models won’t run, won’t provide the right context, or won’t give the right level of services that we’re looking to implement.” Industry surveys consistently cite data quality as one of the greatest barriers to AI success. “The data quality has to be good; otherwise, the user loses confidence in the system,” says Adil.

An effective AI strategy begins with connecting data across the organization and ensuring it is organized, accurate, governed, and accessible in real time. These considerations are most effective when built into models and architecture from the start. Scalable data architecture allows AI systems to evolve alongside the business and connect reliably to the internal information needed to deliver meaningful value.

Gartner predicts that companies will abandon 60% of all AI projects through 2026 if they are not supported by AI-ready data. Avoiding that outcome includes clear data standards and ownership, clean and labeled data, and pipelines that support real-time retrieval.

2. Use context engineering to deliver the right data to every AI query

Context engineering ensures that the model draws on the most pertinent information for each query, selecting and organizing the data needed to produce accurate answers efficiently.

Effective context engineering shapes the inputs that guide AI reasoning and action. While prompt engineering focuses on how a request is worded, context engineering designs the entire information environment around the model: retrieving the right data and presenting it in a structured, machine-readable way. Many organizations are discovering that reliable AI depends as much on context quality as on the strength of the model.

Context engineering relies on a modernized, unified data foundation as well as retrieval and memory systems such as retrieval augmented generation (RAG) and vector databases. It also requires careful prioritization to determine what information matters most, what should be excluded, and when different types of information should be used. Feeding models too much context can dilute relevant details, increase costs, and slow response times.

“Minimum context, correct and current data, and machine-readable information are critical to effective context engineering,” Adil says.

3. Build AI governance and LLM observability in from the start

Strong governance and LLM observability help organizations maintain control over how AI systems use data, monitor system performance, and identify problems before they affect operations.

In the absence of clear controls around retrieval, workflows, and model usage, AI systems often process far more information than necessary. This inefficiency also drives up operating costs by requiring additional computing resources, often reflected in higher token consumption and API charges.

Governance also works in tandem with robust security. AI expands the attack surface, introducing risks such as prompt-based data leakage, model vulnerabilities, and adversarial inputs. Protecting sensitive information requires strong access controls, monitoring, and oversight.

Adil notes that essential controls — including those related to security, granular cost management, project controls, data security, and architecture—are frequently insufficient.

For governance systems to support transparent, compliant, trustworthy, and cost-effective AI, organizations cannot leave them as a layer to add later. Governance structures need to be embedded into architecture, workflows, and decision-making processes from the outset.

When governance is established from the start, it enables robust observability. Observability helps organizations understand how AI applications are performing in practice. Mechanisms for LLM observability and benchmarking allow teams to assess accuracy and utility over time, monitor adoption patterns, and adjust systems as conditions change. Observability also helps organizations gain trust by increasing visibility of model performance, behavior, and failure points.

Furthermore, observability is essential to get ROI of AI initiatives, as the benefits of it are often indirect and business value depends heavily on how systems are adopted and used. Real-time visibility into AI behavior allows organizations to measure performance against expectations, identify gaps between intent and reality, and continuously refine systems as requirements evolve.

In a 2026 report from Elastic, 85% of IT decision makers expect to enable LLM observability for their internal generative AI apps.

“Observability is actually huge. We can use observability data for cost control, decision-making, and engineering efficiency,” Adil says.

4. Keep humans in the loop

The thoughtful design, integration, and governance that maximize AI value demand specialized in-house expertise. Nearly 70% of respondents in Deloitte’s 2025 Tech Executive Survey report plan to grow teams in direct response to generative AI, a clear contrast to widely reported AI-related cuts. Adil agrees: “We think the people aspect is largely what’s going to make AI impactful going forward.”

As AI systems become more embedded in operations, organizations need people who can govern workflows, evaluate outputs, redesign processes, and adapt systems as conditions change. Evolution toward increasingly autonomous tools requires teams skilled in prompt engineering, orchestration, and change management. 

Talent adept at critical thinking and prepared to adapt with technology’s rapid advances will be in high demand. Although turnover brings in fresh thinking, it also presents high costs in system continuity, institutional understanding, and innovation. Human-centered strategy needs to be built into AI execution stages to ensure smooth implementation. 

As Adil says, “Many aspects of the stack are moving very, very fast, but institutional knowledge and the ability to adapt remain durable.

Thoughtful AI investment for future growth

As AI systems evolve from single-task assistants to increasingly autonomous agents, the organizations best positioned to benefit will be those that invest in the underlying systems, governance, and expertise that make AI reliable at scale.

Tech leaders who focus on these fundamentals can move effectively from experimentation to reliable, production-level deployment in the medium term, confident that these elements will remain relevant and adaptable amid constant advancements.

“We fundamentally believe that with these tools, velocity of work will get much faster,” Adil says. “We are really focused on how we can do work with these tools in ways we had not thought of before.”

Learn more about how Elastic is building an AI-first enterprise with these core foundational components.

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.

❌