Enterprise teams building AI agents keep hitting the same wall: a chatbot that can answer a prompt but can't remember what the last five people asked it, and can't tell you whether last month's version actually worked.
In a fireside chat with VentureBeat's Sam Witteveen at VB Transform 2026, Asana's chief product officer, Arnab Bose, unpacked how his team tackled this problem to build a new operating system: Agentic Work Management (AWM). The product treats AI agents as coachable teammates that operate alongside humans rather than as one-to-one assistants.
For product builders and developers trying to move beyond basic integrations, Bose provided a look under the hood. He detailed how Asana engineered AWM, offering a blueprint for solving real-world bottlenecks and building agentic systems at scale.
The Work Graph: 18 years of company data, repurposed
To build an operating system for human-agent teams, Asana needed a ready-made enterprise context graph. They built AWM on top of their 18-year-old architecture: the Work Graph.
This graph-based database organizes information through a structure the company calls the Pyramid of Clarity. The smallest unit of work is a task with an assignee and a due date. Tasks belong to projects, projects roll up into portfolios, and portfolios connect to company-wide goals. The graph can help trace for example how a delayed design task impacts a corporate revenue goal. The Work Graph provides a real-time ledger of who does what, by when, and why.
AWM leverages this architecture to create a multiplayer teammate. A standard AI copilot is stateless and tied to a single user's prompt. Because AWM plugs into the Work Graph, the AI can view overarching company goals, update project statuses, and share memory with human colleagues.
"Because [the agent] is plugged into the Work Graph, it's not just looking at a particular prompt that you're sending it or looking at a particular individual's markdown file system on their local file,” Bose said. “It's working off of that shared ledger for the whole company."
AWM is already in production. Bose said Asana has "several customers live and successful on it," including FedEx, which published its own case study on the shift.
Building in guardrails for confidential work
Shipping AWM to enterprise customers required Asana to solve several technical hurdles. The first was data governance. If an AI teammate acts across a company, it builds a shared memory by learning from workflows and human feedback.
Bose highlighted a critical boundary problem: If an executive uses AWM to build workflows for a confidential project, the system must ensure the agent's updated memory does not leak context to an unauthorized employee who interacts with the same agent later.
"[I] shouldn't be able to leverage that shared memory when I run the AI teammate if you created that memory using that same teammate on a project that is, let's say, a secret M&A project that I don't have access to," Bose said. Asana engineered a system of access controls to govern what triggers the creation of a memory versus the simple execution of a task.
Second, AWM handles dynamic model routing to abstract prompt engineering away from the user. When a user assigns a task to an AI teammate (i.e., drafting a job description for a general manager role), the AI cross-references public job postings, Asana’s internal style guide, and product requirement documents. For a complex task, the system automatically routes the prompt to a heavy frontier model — Bose pointed to Anthropic's Opus and OpenAI's models as examples — while lighter tasks get down-leveled to something faster and cheaper.
"We don't want the knowledge worker to have to think through what the best possible prompt, context engineering, and attachments are that they should put into the task," Bose said. "It should feel as if you were assigning the task to a human being."
This dynamic routing introduces a third challenge: billing abstraction. Agentic tasks vary in computational complexity, making credit burn rates unpredictable.
"We don't want to get into a state where our customers are having to reason about the fact that some of these tasks... are way more complex than others and they'll be burning credits at different rates," Bose said, adding that unpredictable pricing risked customers throttling their own employees by capping how often they could run an AI teammate.
To make AWM commercially viable, Asana designed its billing architecture to charge a static cost per task completion. The platform absorbs the complexity of model selection, token counts, and run limits to ensure predictable enterprise pricing.
The problem with stateless chatbots
AWM targets a specific problem with current enterprise AI deployments: statelessness. Developers can easily connect large language models to enterprise tools like Slack, Google Drive, or Databricks using Model Context Protocol (MCP) integrations. However, basic chat-based agents lack persistence.
Bose detailed a scenario where a user asks a chat agent to draft a marketing campaign based on historical performance and competitive research. The agent fetches data from external tools to answer the prompt, but the execution happens in a vacuum. It is a one-off task that benefits a single individual. It fails to create a reusable workflow for the next person building a similar campaign.
"The challenge with that is that those calls are stateless, and they are not leveraging a shared company brain that is this graph-based database or a context graph," Bose said.
AWM solves this by creating a permanent state. When an AI teammate inside AWM completes a task, the system records the metadata. It registers whether the completion improved the project status and how it moved higher-level company goals.
Inside CoreWeave's product launches
Cloud provider CoreWeave is an early adopter using AWM to overhaul complex new product launches.
"CoreWeave is using both our deterministic AI studio workflow rules as well as multiple AI teammates to do new product launches," Bose shared.
In the past, CoreWeave product managers filled out complicated forms detailing infrastructure, parameters, and costs. Human reviewers manually evaluated these forms and broke them out into specific tasks for finance, marketing, and hardware teams.
Under the AWM workflow, a product manager writes a standard Google document pointing to their product requirement documents. A deterministic AI workflow reads the document, automatically creates the project structure, and assigns tasks. Specialized agents then take over the execution. One agent then watches overall project status and flags bottlenecks; another, working inside individual tasks, forecasts infrastructure costs and recommends approvals when the numbers align with historical budgets. The system automatically triages the busywork while human beings focus on evaluating the AI's outputs.
The frenemy problem
The dynamic gets complicated by the fact that the same frontier-model providers powering AWM under the hood — Anthropic, OpenAI — are also shipping their own competing agent products, like Anthropic's Claude in Slack (Tag). Pressed on the overlap, Bose didn't dispute the tension.
"I think that's the reality that we all have to live in," he said.
His case for AWM's staying power rests on Asana's 18 years of user-experience and workflow data, and prebuilt standard operating procedures for specific industries — expertise he argues raw frontier models don't have. A product like Tag can work well in Slack, he said, but it requires a highly curated channel and its own separate credentials for every downstream app it touches.
"There's a big difference between the power of the model plus a lightweight way to demonstrate its value, and something that's pre-built … for true end-to-end use," Bose said.
At VB Transform 2026, NTT DATA AIVista CEO Bratin Saha joined VentureBeat CEO and editor-in-chief Matt Marshall to discuss the last-mile challenge of operationalizing frontier models in regulated production, where reliability, context, guardrails, and security determine whether AI delivers enterprise value. The conversation centered around the question facing every enterprise now pouring money into AI: how to convert that spending into real, tangible value.
"It's not just a model, you're building a system around the model," Saha said. The last mile is the work of wrapping a frontier model in an enterprise's own data, workflows, and guardrails.
In the end, regulated production turns on more than just technology, Saha said. Today, most enterprise AI projects fail during implementation because of poor integration, domain specialization gaps, lack of governance, and unclear ownership of outcomes. Last-mile specialization turns a capable foundation model into an enterprise agent shaped by domain-specific workflows, risk appetite, client classifications, regulatory interpretations, and institutional knowledge.
Why frontier models stall in enterprise workflows
Frontier models fall well short of production-grade accuracy on many real-world insurance workflows, Saha said, but last-mile specialization can lift them to the reliability enterprises need. Out of the box, those models struggle with the complexity of regulated workflows such as multinational insurance claims.
"These forms are pretty complex, often have handwriting, lots of checkboxes, and so on," he said, and that complexity is why frontier models like Fable 5, Opus 4.8, and GPT-5.5 fall short out of the box.
Saha said the biggest gains come from specializing the entire AI system, not just the foundation model.
That system gets specialized with the customer's data, workflow and, in many cases, the tribal knowledge that never made it into an operating procedure document.
"The biggest bang for the buck comes from the specialization and then these specialized guardrails," he said.
The work has three components:
capturing the enterprise’s context and making it consumable by AI
running an ensemble of models so cost does not go through the roof
and adding specialized guardrails that check the model and force a redo when it gets something wrong.
What the last mile of agentic AI actually requires
None of this involves fine-tuning. VentureBeat’s latest enterprise survey found it ranked last among companies’ model-selection priorities.
Instead, the last mile centers on domain knowledge and undocumented workflows that companies would never expose publicly without losing their competitive edge.
"The last mile is about taking data that's proprietary to you and using that to build a system around the model that can steer the model in the right way that can put the appropriate guardrails around it," Saha said.
In the end, enterprise AI is about moving a workflow from point A to point B rather than deploying a technology, and NTT's advantage comes from pairing AI experts with subject domain experts.
"The only reason is because we go and talk to those human workers and we say, 'How do you actually do the work,'" he said. That expertise is then encoded into an agent.
Success in insurance, manufacturing, and other regulated industries relies on three things at once, he added.
"You need technology, you need the domain expertise, and you need the change management expertise," he explained, adding that across his team's clients, technology is not the bottleneck.
How enterprises turn AI investment into tangible value
For enterprises weighing large AI budgets, Saha's said the payoff comes not from the model but from the work built around it.
"When you're deploying AI in the enterprise, you're not deploying a technology," he said. "You are taking a workflow that exists and taking it from point A to point B." The value is created by the workflow that gets moved, not the model that helps move it.
That reorders where money should go.
"Technology is not the bottleneck," Saha said, pointing instead to the domain expertise and change management wrapped around the model, and to the discipline of commiting to all three together. Spending aimed only at the model leaves most of the return on the table.
Enterprises don’t have to choose between embedding AI into existing workflows and redesigning those workflows from scratch. NTT sees the two as successive stages of the same journey.
"We are starting with embedding in the workflow because it's easier change management," he said, noting that customers running mission-critical operations will not let a vendor rip out a working process midstream. "Once that happens, then we go into, how can we now reimagine this? And that really is where the biggest bang is."
Where enterprise AI stays bespoke and where it becomes scalable
Keeping intelligence in the surrounding system rather than the model also preserves swappability and lets enterprises take advantage of open-weight and open-source models as they mature. Saha’s team runs an ensemble that mixes frontier and open-source models, and he expects the industry to lean on open weights wherever the cost of a mistake is low while reserving frontier reasoning for the cases that demand it.
"In many situations, especially in regulated industries where mistakes are very expensive, that last extra couple of percent matters," he said.
The platform follows the same pattern: Guardrail generation and neurosymbolic models scale across customers, while capturing each organization’s tribal knowledge remains bespoke. Saha pointed to NTT DATA’s position as one of the world’s largest insurance third-party administrators as an advantage in acquiring that expertise.
"The ability to take that knowledge and trust that has been built over 20 years is very hard to replicate instantly, and I do think that is a durable aspect of what we have," he said.
Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
If you have built anything with retrieval-augmented generation (RAG) in the last two years, you have lived its central frustration: You chop your documents into chunks, embed them, retrieve the top few that look similar to the question, and hand them to the model. For “What was our Q3 refund policy?” This works beautifully. For “What are the recurring themes across two years of customer complaints?” it falls flat — because no single chunk contains the answer.
The fashionable fix is GraphRAG: Instead of feeding the model isolated snippets, you first build a knowledge graph of the entities and relationships in your corpus, then use that structure as context. The pitch is seductive. But seductive pitches deserve scrutiny, so I went through the evidence — the original Microsoft paper plus four independent benchmark studies — to answer a simple question: When you swap text chunks for a context graph, do answers actually get better?
The short version: Yes, substantially — but only for the right kind of question, and not for free. Let me show you the receipts.
Why text chunks hit a wall
Standard vector RAG retrieves the k passages most similar to your query. That design has three structural blind spots:
It can’t connect the dots. When an answer requires joining facts that live in different passages through a shared entity, chunks embedded in isolation never reveal the link.
It’s blind to global questions. “What are the main themes?” needs the whole corpus, but similarity search only returns the handful of chunks that superficially resemble the question.
It severs context at chunk boundaries. The relationships and hierarchy that complex reasoning depends on are exactly what chunking throws away.
Microsoft Research framed this crisply when they introduced GraphRAG: Baseline RAG “struggles to connect the dots” and performs poorly when asked to “holistically understand summarized semantic concepts over large data collections.”
What a context graph changes
GraphRAG attacks the problem before any question is asked. During indexing, a large language model (LLM) reads every chunk and extracts entities, relationships, and claims, assembling them into a weighted knowledge graph. It then runs community detection (the Leiden algorithm) to cluster the graph into a hierarchy of related topics, and pre-writes a natural-language summary for each community.
At query time, those summaries do the heavy lifting. Each relevant community drafts a partial answer (the “map” step), the partials are ranked and merged (the “reduce” step), and the model synthesizes a final response grounded in structure rather than in a few cherry-picked snippets. Variants like HippoRAG take a different route, using the graph plus a Personalized PageRank walk to find the right passages — but the core idea is the same: Let relationships, not just cosine similarity, decide what context the model sees.
The evidence: Four studies, one pattern
1. Global sense making: The headline win
Microsoft pitted GraphRAG head-to-head against naïve RAG on global, “make sense of the whole corpus” questions over million-token datasets, with an LLM acting as judge across three axes: Comprehensiveness, diversity, and empowerment.
GraphRAG won 72 to 83% of comprehensiveness comparisons and 62 to 82% of diversity comparisons against vector RAG. Its highest-level summaries used up to 97% fewer tokens than processing the source text directly.
That is not a rounding-error improvement. On exactly the kind of question that breaks text-chunk RAG, the graph wins two out of three times or better.
2. Multi-hop retrieval: The graph finds what chunks miss
The second piece of evidence is about retrieval quality: Does the right supporting passage even make it into the top results? On the standard multi-hop QA benchmarks (MuSiQue, HotpotQA, 2WikiMultiHopQA), graph-guided retrieval lifts Recall@5 dramatically:
Average Recall@5 climbs from 73.4% (naïve RAG) to 87.8% (graph-guided), a +19.6 point gain.
The biggest jumps come on the hardest, cross-document sets: +31 points on MuSiQue and +28 points on 2Wiki.
HippoRAG reports up to a 20% accuracy improvement on multi-hop QA, at 10–20× lower cost and 6–13× faster than iterative retrieval methods.
3. The controlled head-to-head - where it gets honest
Here is where the story gains nuance. A 2025 study from Michigan State and Meta ran RAG against four GraphRAG families under one unified protocol — identical chunking, embeddings, and generation — and found no single winner. The two approaches are complementary:
On single-hop, factual lookup (natural questions), plain RAG edged ahead (F1 64.8 vs. 63.0 for the best graph method).
On multi-hop reasoning (MultiHop-RAG), graph-guided retrieval pulled in front (70.3 vs. 67.0 overall accuracy).
The lesson: A context graph is not a universal upgrade. It is a specialized one that pays off precisely when questions demand reasoning across pieces.
4. When to use graphs: The task-type verdict
The most recent benchmark, GraphRAG-Bench (ICLR 2026), set out to answer “In which scenarios do graph structures provide measurable benefits?” Its accuracy-by-task numbers map the boundary cleanly:
Simple fact retrieval: Text chunks 60.9 vs. graph 60.1 — effectively a tie. The graph’s structure is overhead the query doesn’t need.
Complex reasoning: Graph 53.4 vs. chunks 42.9 — a +10 point graph win.
Contextual summarization: Graph 64.4 vs. chunks 51.3 — a +13 point graph win.
The scorecard
Read top to bottom, the pattern is unmistakable: The graph’s advantage grows with the reasoning depth of the question, while text chunks hold their ground on isolated facts.
The catch: Cost and the LLM-judge problem
Two caveats keep this from being a slam dunk, and ignoring them is how teams end up disappointed.
Building the graph is expensive. Having an LLM extract entities and relationships from an entire corpus isn’t cheap. One analysis put index construction at roughly $48 against GPT-4o for a moderate corpus, far above a vanilla vector index. (Microsoft’s own follow-up, LazyGraphRAG, defers extraction to query time and cuts that to around 0.1% of the cost - a tacit admission that the original budget is impractical for many deployments.)
Many of the wins are judged by another LLM — and LLM judges are biased. An independent audit found systematic flaws in this evaluation style: position bias (swapping which answer appears first can swing the win-rate by more than 30 points), length bias, and trial bias (identical comparisons disagree across runs). After correction, one popular method’s reported 66.7% win rate fell to about 39% — below the 50% break-even line.
The takeaway is not “the research is wrong.” It is that the large gains — the +20% multi-hop accuracy, the +15-to-30-point recall jumps — are robust, while narrow comprehensiveness margins deserve a skeptical second look with reference-based metrics.
So when should you reach for a context graph?
Strip away the hype and the decision is refreshingly practical.
Use a context graph when: Your questions are multi-hop, global, or sensemaking in nature; you need comprehensive, multi-perspective answers; and your corpus is richly interconnected (research libraries, case files, incident histories, knowledge bases).
Stick with text chunks when: Your queries are mostly single-fact lookups; your corpus is small or flat; and indexing cost, latency, and operational simplicity outweigh a marginal quality bump.
Best of all, go hybrid: The systematic studies converge on the same recommendation: route each query to the right method, or fuse evidence from both. Combining graph and chunk retrieval consistently beats either one alone. You don’t have to choose a religion; you have to build a router.
The bottom line
A context graph is not magic, and it is not snake oil. It is a targeted instrument. Hand it a question that requires connecting scattered facts or synthesizing a whole corpus, and it will outperform text chunks decisively. Hand it “what’s the phone number on page 3,” and you’ve paid for indexing you didn’t need.
The teams that win with GraphRAG in 2026 won’t be the ones who graph everything. They’ll be the ones who know which questions deserve a graph — and build pipelines smart enough to tell the difference.
Dattaraj Rao is an R&D architect at Persistent Systems
If you ask an AI coding agent to write a standalone Python script to parse a single JSON file, it will likely give you a perfect answer in seconds. But the same agent often breaks if you ask it to build a systematic data processing pipeline, like ingesting thousands of messy documents, chunking text, scoring quality, and filtering noise for a Retrieval-Augmented Generation (RAG) system that fits your specific enterprise stack.
While large language models (LLMs) excel at one-off code generation, their outputs for complex data-processing tasks are typically free-form, disposable scripts. These scripts are detached from the governable workflow abstractions that MLOps teams rely on for production, making them difficult to audit or edit visually.
To address this, researchers at Peking University, Zhongguancun Academy, and Shanghai’s Institute for Advanced Algorithms Research introduced DataFlow-Harness, an open-source framework that guides an LLM agent to build structured, visual data-processing workflows step-by-step, rather than writing raw code from scratch.
The framework makes AI-generated pipelines easier to manage and integrate into existing architectures because the generated artifacts are persistent and easily editable.
The researchers report that the platform achieves a 93.3% observed end-to-end pass rate on a 12-task data-engineering benchmark. Compared to standard Claude Code, it reduces API costs by up to 72.5% and response latency by 49.9%, while achieving nearly the same success rate as an AI given the entire codebase to write standard scripts. For enterprise teams, this means getting the speed of AI automation without accumulating unmanageable technical debt, ensuring that pipelines remain secure, auditable, and ready for production.
The "NL2Pipeline gap"
Data-centric AI requires workflows for tasks like synthetic data generation, retrieval augmentation, and model training. While LLMs can translate natural language into executable implementations to perform these tasks, high task accuracy is insufficient for production deployment.
"The first wall is usually not writing Python," Runming He, first author of the DataFlow-Harness paper, told VentureBeat. "Modern coding agents can often produce a plausible script quickly. The harder problem is grounding that script in a live production platform: using operators that are actually installed, matching the real dataset schema, referring to registered datasets and model services, preserving dependencies between stages, and leaving behind an artifact that another engineer can understand and revise."
General-purpose AI agents frequently hallucinate dependencies, relying on unavailable operators or outdated platform assumptions. Instead of leaving behind an artifact that another engineer can understand and revise, they generate disposable code that is difficult to audit through workflow managing tools.
The researchers define this challenge as the "NL2Pipeline gap": the disconnect between a user expressing workflow requirements in natural language and the production environment requiring structured and persistent pipeline assets.
The researchers demonstrated this gap in their experiments. For example, when Claude Code was allowed to write standard, free-form scripts using codebase context, it hit a 94.2% success rate. However, when restricted to only using the platform's specific building blocks to create a native workflow graph, its success rate dropped to 83.3%. This gap is the paper's central finding: native, governable pipelines are meaningfully harder for the agent to produce than throwaway code.
“Closing this gap requires more than improving code-generation accuracy: construction must remain grounded in platform semantics and produce artifacts that integrate with the host platform,” the researchers write.
How the four components work together
"DataFlow-Harness changes the agent’s action space," He said. "Instead of asking the agent to emit arbitrary code, it retrieves the live operator registry and current pipeline state through MCP and applies typed, incremental changes to a persistent DAG."
To achieve this, the platform organizes workflow synthesis around four components: the Data Pipeline Backend, the interaction layer (DataFlow-WebUI), the MCP Tools Layer, and the AI guidance layer (DataFlow-Skills).
The Data Pipeline Backend acts as the authoritative source of truth across conversational, visual, and programmatic interfaces. It represents the pipeline as a directed acyclic graph (DAG), a structured workflow map containing data sources, configured pre-built processing modules (which the researchers refer to as "operators"), and execution dependencies. Instead of generating free-form code, agents interact with this backend through “typed mutations,” like adding an operator or connecting edges.
DataFlow-Skills are markdown files that inject domain-specific knowledge into the model's context window, guiding it on operator-selection patterns, schema inference, and assembly procedures. Rather than letting the AI guess how to assemble components, skills provide the AI with compatibility rules, teaching it how to correctly match different data formats and handle complex data structures without breaking the pipeline.
The MCP Tools Layer gives the AI access to the operator registry and current state of the data workflow. The AI proposes structured changes through the tools layer. The system validates the changes to ensure the workflow runs in a valid sequence and that every connected module speaks the same data language.
DataFlow-WebUI provides two interfaces that allow humans and AI to build the workflow together. Developers can describe workflow requirements in natural language through a conversational interface. They can also access the workflow as a graphical map in a visual DAG editor. Here, they can directly inspect the changes proposed by the AI and make modifications.
“The current implementation performs static checks against platform metadata before accepting pipeline changes,” He said. “These include checks for registered datasets, operators and model-serving references, field flow, and some invalid parameter usage, as well as structural validity. The result is visible in a graphical editor and can be revised either manually or by the agent in later turns.”
The results: 93.3% pass rate, 72.5% lower cost
The researchers tested DataFlow-Harness on a benchmark of 12 tasks across six industrial data-processing scenarios, such as QA generation, review governance, and schema normalization. They used Claude Opus 4.7 as the backbone model in their experiments.
They compared DataFlow-Harness against three baselines:
Vanilla CC: An unconstrained coding baseline using standard Claude Code.
Context-Aware CC: An agent that has access to the DataFlow codebase in its context window.
MCP-only: An agent that has access to the DataFlow MCP tools and is instructed to generate platform-native DAGs (without access to DataFlow-Skills).
DataFlow-Harness achieved a 93.3% end-to-end pass rate, improving by 10.0 percentage points over MCP-only and beating Vanilla CC (91.7%), while being within 0.9 percentage points of Context-Aware CC (94.2%).
Importantly, it reduced API costs to $0.261 per task, a 72.5% drop compared to Vanilla CC and 42.8% compared to Context-Aware CC. In generating workflows, it was 49.9% faster than Vanilla CC and 17.6% faster than Context-Aware CC.
DataFlow-Harness proved particularly effective on complex tasks that depend on implicit domain knowledge, like QA generation. The baseline MCP-only approach frequently generated structurally valid DAGs but struggled to infer task-specific procedures from operator descriptions alone.
To show how this works in the real world, the researchers detailed a textbook-to-VQA extraction task. This job required the AI to stitch together capabilities such as PDF parsing, layout recovery, OCR, figure extraction, multimodal understanding, and long-range question-answer matching. DataFlow-Harness achieved 97.2% precision and an 87.3% coverage rate, easily beating the baselines. By having the AI snap together existing platform assets rather than coding complex tasks from scratch, it recovered more valid QA pairs from the document.
Their experiments also showed that DataFlow-Harness is highly effective at creating data generation pipelines. For example, in a synthetic instruction-data generation task, the agent built a multi-stage pipeline that generated candidate instruction–response pairs, critiqued and rewrote them, scored them with an LLM-based judge, and filtered low-quality outputs before training.
"Such workflows are costly to build and fragile to maintain as collections of ad hoc scripts," He said. "The harness does not make them automatically safe, but it turns them into explicit, editable stages that engineers can inspect, test, and govern using normal production controls."
Similarly, when tasked with building a math data cleaning-and-synthesis pipeline, the data produced by the DataFlow-Harness pipeline trained a better-performing model with higher average accuracy on AIME24 and AIME25 benchmarks than the data produced by the vanilla Claude Code pipeline.
Tech stack fit and implementation tradeoffs
For engineering teams evaluating DataFlow-Harness, it is important to understand how it fits into existing infrastructure. Released under the Apache 2.0 license, the current implementation requires a bit of engineering to fit into popular tech stacks.
"The current implementation is native to the DataFlow platform; it is not a turnkey Airflow, Prefect, or Spark plug-in," He said. To use those systems as an execution backbone, teams must build an adapter to connect their organization’s registry, metadata, and execution interfaces to the agent's control layer.
Furthermore, organizations must invest in the boundaries they want the AI to respect. This requires maintaining an operator registry, defining schemas, and encoding recurring domain procedures as Skills. Because of this overhead, He recommends against using the framework for small, one-off transformations where a simple script suffices, or in legacy environments that cannot expose reliable metadata.
Finally, while the platform prevents illogical connections by validating structural properties, it is an engineering control layer, not a compliance substitute. "The harness should still be treated as an engineering control layer, not as a substitute for compliance policy, validated detection models, access controls, audit logging, or human approval," He said.
The platform is open-source, and developers can access the source code and codebase documentation directly via the project's GitHub repository.
As protocols like MCP become standardized, the boundary between human engineers and AI agents will shift. "The goal is not autonomous data engineering without oversight," He said. "It is a better division of labor: agents perform repetitive construction inside explicit boundaries, while engineers remain responsible for the semantics, policies, and consequential decisions that require domain accountability."
The AI agent observability space is taking off — but how can enterprises be sure what observability products and solutions they need?
Observability startup groudcover (lower case "g" intentional) announced this week that it raised $100 million in a round led by One Peak, bringing its total funding to $160 million.
The company says it has more than 250 paying customers, tripled annual recurring revenue over the past year and is increasingly replacing established observability platforms inside enterprise environments. Those are company-reported figures, but together they point to growing momentum in one of enterprise software's most competitive markets.
That market has long been dominated by companies including Datadog, Dynatrace, New Relic, Splunk and Grafana. Between them, they represent billions of dollars in annual revenue and years of product maturity. Breaking into that group has never been easy.
groundcover's argument is that artificial intelligence has fundamentally changed the assumptions those platforms were built on.
Rather than competing feature for feature, the four-year-old company is trying to convince enterprises that the architecture underpinning observability itself needs to change as AI systems become more autonomous, produce vastly more telemetry and increasingly participate in software operations. Whether that thesis proves correct remains an open question, but it offers a compelling lens through which to examine how observability is evolving alongside enterprise AI.
AI is turning telemetry into an infrastructure problem
Observability has traditionally been viewed as a post-production discipline. Engineers deploy applications, monitor logs, metrics and traces, investigate incidents, and improve reliability over time.
That workflow is changing.
AI-assisted software development has dramatically accelerated deployment cycles. Coding assistants generate more code, infrastructure evolves more rapidly, and organizations are deploying increasingly complex distributed systems that combine microservices, Kubernetes clusters, APIs and large language models. At the same time, enterprises are beginning to operate AI agents that execute multi-step workflows, call external tools and interact with production systems.
Each of those activities generates telemetry.
The result is an explosion of operational data that organizations increasingly want to retain rather than discard. AI applications introduce additional layers of observability beyond traditional infrastructure monitoring, including prompt execution, model latency, token consumption, retrieval pipelines, tool invocations and agent behavior. As enterprises experiment with autonomous systems, that telemetry becomes increasingly valuable because it provides the context needed to understand what an AI system actually did and why.
For many organizations, this creates tension with pricing models that charge according to the amount of data ingested.
Historically, engineers have often responded by sampling traces, shortening retention periods or limiting which data is collected. Those approaches reduce costs, but they also reduce visibility precisely when AI-driven systems demand more complete operational context.
"We've seen telemetry exploding," groundcover co-founder and CEO Shahar Azulay said during a recent media briefing. "Users are frustrated by not getting all the value from Datadog and similar platforms. They're limiting the data, siloing it, sampling it."
Whether that frustration is widespread enough to reshape the market remains to be seen, but the underlying trend is difficult to ignore. AI is making observability less about collecting enough data and more about collecting everything organizations may eventually need.
Rather than adding AI, groundcover argues the architecture itself has to change
Many observability vendors have introduced AI assistants, AI-powered root cause analysis and AI observability features over the past two years. Datadog, Dynatrace, New Relic and Grafana have all announced products aimed at helping enterprises monitor AI applications or automate operational tasks.
groundcover acknowledges those developments but argues they do not address what it sees as the more fundamental issue: where telemetry lives and how customers pay for it.
Instead of operating a conventional SaaS platform that stores customer telemetry in vendor-managed infrastructure, groundcover uses what it calls a bring-your-own-cloud (BYOC) architecture.
Customers keep the data plane—including telemetry storage and processing—inside their own AWS, Microsoft Azure or Google Cloud environments, while groundcover provides a managed control plane and user experience. A fully self-hosted deployment option is also available.
While some competitors, including Datadog and a few other observability vendors, do offer limited hybrid or customer-controlled data residency options, these are generally not equivalent to a full BYOC model. In most cases, telemetry is still processed and stored within the vendor’s managed infrastructure, with only partial controls (such as regional data residency, private links, or selective log forwarding) available.
That architectural decision influences nearly every aspect of the company's strategy.
Because customers already pay for their own cloud infrastructure, groundcover argues it can avoid charging based on telemetry ingestion. Instead, pricing is based primarily on monitored hosts, regardless of telemetry volume.
The company believes this changes customer behavior.
Rather than deciding which logs or traces are too expensive to keep, organizations can theoretically retain complete telemetry and use it for operational analysis, compliance and AI-assisted troubleshooting.
"We don't price by data volume," Azulay said. "We price by the size of the infrastructure."
The distinction matters because AI workloads tend to increase telemetry far faster than infrastructure itself.
That does not necessarily make host-based pricing universally cheaper. Organizations with relatively light workloads spread across many hosts may find different economics than dense Kubernetes environments generating enormous amounts of telemetry. The company's own briefing notes that per-host pricing is most advantageous for organizations with high telemetry density and may be less compelling for lightly utilized fleets.
Still, the broader argument is less about cost alone than predictability. Enterprise infrastructure teams often struggle with observability bills that fluctuate alongside application growth. groundcover's model attempts to align pricing more closely with infrastructure planning rather than data generation.
eBPF sits at the center of the company's technical differentiation
The second pillar of groundcover's strategy is eBPF, a Linux kernel technology that has rapidly become one of the most important building blocks for modern cloud observability.
Instead of requiring developers to manually instrument applications, eBPF allows software running inside the operating system kernel to observe network traffic, system calls and application behavior with minimal code changes.
That enables faster deployment and broader visibility across infrastructure.
For organizations operating Kubernetes clusters and cloud-native applications, reducing instrumentation complexity can significantly shorten deployment times while increasing telemetry coverage.
Azulay argues this becomes especially important as AI systems generate increasingly complex interactions across services.
"Our sensor allows us to observe systems very deeply from infrastructure to application to AI workloads without developers needing to instrument code," he said during the briefing.
eBPF itself is hardly unique. Many observability vendors now incorporate it into their platforms.
What groundcover argues differentiates its approach is combining automatic eBPF collection with customer-controlled storage, OpenTelemetry compatibility and unified pricing inside a single platform.
The company's own research briefing acknowledges that none of these technologies individually represents a competitive moat. The claimed differentiation lies in the combination of eBPF-first collection, managed BYOC architecture, host-based economics and full-stack observability delivered together.
AI agents are becoming both customers—and users—of observability
Perhaps the most interesting aspect of groundcover's strategy extends beyond traditional monitoring.
The company increasingly describes observability as infrastructure for autonomous software development.
Historically, observability platforms have served human operators investigating production incidents.
groundcover believes future observability platforms will increasingly serve AI agents as well.
Its Agent Mode product allows engineers to investigate incidents using natural language across logs, metrics, traces and Kubernetes events. More importantly, Azulay envisions observability becoming the feedback mechanism that informs coding agents about what actually happened in production.
Rather than simply detecting failures after deployment, observability becomes continuous operational context that autonomous systems can use to evaluate changes, identify regressions and eventually recommend or implement fixes.
"We're seeing observability moving from being a post-production tool... to people taking context from production and feeding it back to their coding agents so they can write code better," Azulay said.
Today, the company emphasizes that humans remain in the loop.
Agent Mode investigates incidents and surfaces recommendations, but production changes still require human approval. Azulay expects autonomy to increase gradually as organizations become more comfortable allowing AI systems to participate in operational workflows.
That vision reflects a broader trend emerging across enterprise software, where AI agents increasingly span development, testing, deployment and operations rather than functioning as isolated assistants.
Why some enterprises are considering alternatives
groundcover is entering an intensely competitive market populated by vendors with decades of enterprise experience.
Datadog alone generated more than $3 billion in annual revenue in 2025. Dynatrace, Cisco's Splunk business, Grafana Labs and New Relic all maintain extensive partner ecosystems, mature integrations and enterprise support organizations that newer entrants cannot easily replicate.
groundcover is not attempting to outscale those incumbents overnight.
Instead, it argues that AI creates an architectural inflection point similar to previous transitions from on-premises infrastructure to cloud-native computing.
According to Azulay, many customers initially adopt groundcover to reduce observability costs but increasingly remain because they want unrestricted access to richer telemetry and AI-native workflows.
He says deployments typically replace incumbent platforms rather than operate alongside them, although the company has not publicly disclosed customer migration data or independent studies validating that claim.
The company's journalist briefing also urges caution around some performance claims.
Revenue growth, customer counts and enterprise adoption figures originate from groundcover itself. Published customer case studies reporting significant cost savings are vendor-authored and should not be treated as independent validation without additional evidence. The briefing also recommends scrutinizing exactly what metadata leaves customer environments in standard BYOC deployments, rather than assuming that no operational data ever reaches vendor infrastructure.
Those caveats are important because the observability market has become crowded. Gartner currently tracks more than one hundred observability products, and nearly every major vendor now markets AI-powered operational capabilities.
Success will likely depend less on whether AI matters—which increasingly appears inevitable—and more on whether enterprises conclude that existing architectures remain sufficient.
The larger question investors are betting on
Viewed narrowly, groundcover's Series C is another large infrastructure funding round.
Viewed more broadly, it reflects a growing debate about what observability becomes in an era where software increasingly writes, tests and operates itself.
If AI continues generating exponentially larger volumes of operational data, traditional assumptions about telemetry collection, pricing and storage may come under increasing pressure. Vendors that built businesses around charging for data ingestion may need to evolve their economics alongside customer expectations. New entrants, meanwhile, have an opportunity to design around those changing assumptions from the outset.
groundcover believes that opportunity lies in combining customer-controlled infrastructure, automatic telemetry collection and AI-assisted operations into a platform designed for autonomous software rather than simply adding AI features to existing observability products.
Whether that architectural bet proves durable will depend on enterprise adoption over the next several years.
But the company's latest funding round suggests at least some investors believe the next battle in observability will not be fought over dashboards or alerts. It will be fought over who builds the operational data layer that increasingly intelligent software relies upon to understand—and eventually manage—the systems it runs.
Days after OpenAI disclosed that two frontier AI models escaped containment measures and autonomously cyberattacked the AI code sharing platform Hugging Face, OpenAI's top U.S. rival Anthropic tonight revealed that — lo and behold — it has also had models surreptitiously access the web when they weren't supposed to, and cyberattack and gain "unauthorized access" to three other organizations.
Anthropic says that it ran "capture the flag" cybersecurity scenarios with three models — Claude Opus 4.7, Claude Mythos 5, and unnamed internal research prototype — with its partner, the AI security firm Irregular. Anthropic says the models were not supposed to have internet access, but that a misunderstanding with Irregular allowed them to access the internet. Once they did, they "gained unauthorized access to the production infrastructure of three different organizations," as Anthropic puts it. Anthropic's blog post continues:
"Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.
However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment."
Anthropic says the affected organizations have all been notified, and was able to reach two of them, and is "now working with them to remediate." The third hasn't yet been reached.
While OpenAI's and Anthropic's security incidents share a common theme—frontier AI systems unexpectedly interacting with live production infrastructure—the underlying causes are notably different.
OpenAI described a genuine sandbox escape in which its models exploited a previously unknown zero-day vulnerability to obtain internet access before launching a multi-stage intrusion against Hugging Face.
Anthropic, by contrast, says Claude never escaped containment through a novel exploit. Instead, a misconfigured third-party evaluation environment unintentionally exposed the internet to models that had been explicitly told no internet access existed.
The distinction matters. Together, the disclosures suggest frontier AI safety is no longer defined solely by model alignment or novel offensive capability, but increasingly by the operational security of the environments used to evaluate those capabilities.
What Anthropic found
Anthropic said it reviewed 141,006 cybersecurity evaluation runs after OpenAI published its Hugging Face report. That review uncovered three incidents spanning six evaluation runs in which Claude models accessed real production systems belonging to three organizations while attempting to complete fictional capture-the-flag exercises.
The incidents involved Claude Opus 4.7, Claude Mythos 5 and an unreleased internal research model.
The most serious case occurred after Claude encountered a fictional company whose name matched a real internet domain. Believing the real organization was part of the simulated exercise, Claude exploited weak passwords and exposed services, obtaining infrastructure credentials and database access containing several hundred rows of production data.
A second incident was arguably even more unusual. Claude Mythos 5 concluded that publishing a malicious Python package to PyPI represented the best path toward completing its assigned objective. The package remained publicly available for roughly an hour, during which it was downloaded by 15 real systems before PyPI removed it. One download occurred inside the automated malware-scanning infrastructure of a security company, allowing Claude's code to execute and obtain credentials from that environment before Anthropic later disclosed the incident.
The third incident involved an internal research model that scanned approximately 9,000 internet-facing systems before compromising one organization using well-known techniques including exposed debug credentials and SQL injection. Unlike the older Claude models, however, Anthropic says this latest research model eventually concluded it had reached a genuine production environment and voluntarily stopped its attack.
Similar outcomes, different root causes
Although the headlines inevitably invite comparison with OpenAI's disclosure, the two incidents illustrate different failure modes.
OpenAI's investigation concluded that its models actively discovered and exploited a previously unknown vulnerability in an Artifactory package-cache proxy to escape an otherwise isolated evaluation environment before later compromising Hugging Face through additional chained vulnerabilities.
The company characterized that event as an "unprecedented cyber incident" involving state-of-the-art offensive AI capabilities.
Anthropic's report instead argues its models never needed to break containment because containment had already been weakened through operational misconfiguration. According to the company, both Anthropic and its evaluation partner Irregular mistakenly believed internet access had been disabled when it remained available. Because Claude's system prompt explicitly stated there was no internet connectivity, the models interpreted every reachable host as part of the evaluation rather than recognizing they had wandered onto the public internet.
Anthropic therefore characterizes the incidents primarily as infrastructure and evaluation-harness failures rather than evidence of models independently pursuing unauthorized goals.
The reports nevertheless converge on one uncomfortable conclusion: frontier AI systems are increasingly capable of executing long-horizon offensive cyber operations whenever evaluation environments permit them to do so.
Four major enterprise security takeaways so far...
For enterprise security leaders, Anthropic's disclosure arguably shifts the conversation beyond "Can frontier models escape?" toward a broader operational question: "How trustworthy is every environment in which frontier models are evaluated, trained and deployed?" There are at least 4 lessons to be learned:
The first lesson is that evaluation infrastructure itself now deserves production-grade security engineering. Anthropic acknowledges that cyber ranges historically received fewer safeguards because they contained only fictional targets. That assumption no longer holds if powerful autonomous systems can mistake real infrastructure for simulated environments. Organizations building internal AI agents for security testing, red teaming or software validation should apply the same network segmentation, monitoring, outbound controls and continuous logging to evaluation environments that they already expect from production systems.
Second, both disclosures reinforce that alignment alone cannot compensate for environmental ambiguity. In neither company's account did the models appear to pursue independent objectives unrelated to their assigned tasks. Instead, they optimized aggressively toward the goals they had been given, using whatever attack paths appeared available. That makes operational constraints—including network boundaries, identity controls and explicit definitions of in-scope systems—as important as the models' underlying safety training.
Third, enterprises deploying increasingly autonomous AI agents should treat situational awareness as a security dependency rather than an academic capability. Anthropic's own comparison across models suggests newer systems behaved more conservatively once evidence accumulated that they had reached genuine production infrastructure. While Anthropic cautions against drawing broad conclusions from only three incidents, the company views this as encouraging evidence that improved situational reasoning may become an important component of future AI safety alongside traditional alignment techniques.
Finally, these two disclosures together mark an inflection point for enterprise threat modeling. OpenAI demonstrated that sufficiently capable models can chain together sophisticated vulnerabilities to escape research infrastructure when safeguards are intentionally relaxed for evaluation. Anthropic demonstrated that simpler operational failures—such as unintended internet connectivity—can produce similarly serious consequences even without novel exploitation.
The common denominator is not any single vendor or model family. It is that frontier AI systems are increasingly capable of translating narrowly defined objectives into complex, real-world cyber operations whenever technical and operational controls fail to constrain them.
For enterprise CISOs, that means AI safety can no longer be viewed solely as a model problem. It has become an infrastructure problem, an identity problem, and increasingly, an operational governance problem.
Inkling Small is a 276-billion-parameter multimodal reasoning model with a permissive Apache 2.0 license that comes within a single point of its larger sibling on the third-party Artificial Analysis Intelligence Index, despite the original Inkling being 975 billion parameters (internal model settings). It accepts text, image and audio inputs, produces text, and supports a context window of up to one million tokens.
Inkling Small uses 12 billion active parameters per token, compared with Inkling’s 41 billion active parameters, while preserving much of the flagship’s coding, reasoning and multimodal performance.
For enterprises, the appeal is not simply that Inkling-Small is smaller. It is that developers appear to give up relatively little capability while reducing the model’s compute requirements, inference costs and deployment footprint.
The model remains far too large for a laptop or conventional workstation, but it is materially easier to operate than the 3.5X larger flagship, making it a good fit for enterprises with some — but not a lot — of their own graphics processing units (GPUs).
Thinking Machines has released the full weights on Hugging Face and added support for fine-tuning through its Tinker model training application programming interface (API).
At launch, the company is advertising a limited-time 50% discount, bringing API pricing for the standard 64K-context Inkling-Small model to $0.58 per million prefill (input) tokens, $1.44 per million sampled (output) tokens, and $1.73 per million training tokens, with cached prefill requests priced at $0.116 per million tokens. A 256K-context variant is also available at higher rates.
Nearly the same performance at a quarter the size
Artificial Analysis assigned Inkling-Small a score of 40 on its Intelligence Index, compared with 41 for Inkling.
That result is notable because Inkling-Small has 276 billion total parameters and 12 billion active parameters, while Inkling has 975 billion total parameters and 41 billion active parameters.
Artificial Analysis also reported that no open-weight model at Inkling-Small’s size or smaller scored higher on the index.
The model does more than merely approach the flagship’s aggregate score. On several evaluations, it surpasses Inkling.
Thinking Machines reports that Inkling-Small scores 80.2% on SWE-bench Verified, compared with Inkling’s 77.6%, and 64.7% on Terminal Bench 2.1, compared with 63.8% for the larger model. It also edges ahead on SciCode, Humanity’s Last Exam, GPQA Diamond and CritPt.
The gains are not universal. Inkling retains a clear advantage on factual knowledge and some agentic tasks. Inkling-Small scores 15.5% on τ³-Banking, compared with 23.7% for Inkling, and its AA Omniscience score is negative, reflecting weaker factual coverage even though its reported hallucination rate is slightly lower.
That tradeoff matters for enterprises. Inkling-Small may be attractive for coding assistants, tool-use systems, retrieval-augmented generation, document analysis and multimodal workflows, but organizations using it for high-stakes factual tasks will still need retrieval, verification and human review.
How a 276B model uses only 12B parameters at a time
Inkling-Small is a sparse Mixture-of-Experts model. According to the model card published by Thinking Machines, its 42-layer decoder routes each token to six of 256 specialized experts, along with two shared experts that remain active for every token.
That architecture helps explain the distinction between the model’s 276 billion total parameters and its 12 billion active parameters. The system retains a large pool of learned capacity but activates only a fraction of it during each inference step.
It is also natively multimodal. Images, audio and text are projected into a shared representation and processed jointly by the decoder rather than being handled through completely separate external systems. Thinking Machines lists coding assistants, agentic applications, chatbots, RAG systems and other multimodal applications among its intended uses.
The company also supports variable reasoning effort, allowing developers to increase or reduce the model’s test-time compute depending on the difficulty of the task. That gives engineering teams a direct way to balance quality, latency and cost across different workloads.
Unfortunately, small does not mean it runs on a laptop
Despite its name, Inkling-Small is not a consumer-scale model.
The standard BF16 checkpoint requires at least 600 GB of aggregate GPU memory, according to Thinking Machines. The company lists two supported configurations: 4x NVIDIA B300 GPUs or 8x NVIDIA H200 GPUs.
A quantized NVFP4 checkpoint lowers the requirement to roughly 180 GB of aggregate VRAM. Thinking Machines says that version can run in W4A4 mode on a single NVIDIA B300, or in W4A16 mode on two H200 GPUs.
That rules out ordinary laptops, MacBooks, desktop gaming PCs and most developer workstations. Even heavily equipped local systems generally fall far short of the required memory.
The practical deployment targets are enterprise GPU servers, cloud clusters and specialized inference providers. The “Small” label is therefore relative to Inkling, not to the broader universe of local models.
Still, the reduction is meaningful. A model that approaches Inkling’s performance while needing substantially less aggregate memory can lower hosting costs, make capacity planning easier and widen the group of organizations capable of self-hosting it.
For companies that want control over data, model behavior and fine-tuning, that smaller footprint may be more important than chasing the highest possible benchmark score.
And of course, it being open source means that it will no doubt be rapidly quantized (made less precise but requiring less compute) and likely blended with other models to be made even smaller for consumer-grade hardware.
Apache 2.0 is the gold standard for enterprise open source models
The licensing may be as important as the benchmarks.
Inkling-Small is released under Apache 2.0, one of the software industry’s most familiar permissive licenses. It generally allows organizations to use, modify, fine-tune, redistribute and commercialize the model, including inside proprietary products, provided they comply with the license’s notice and attribution requirements.
That gives enterprises far more legal flexibility than many custom “open” AI licenses, which may include revenue thresholds, branding obligations, use restrictions or separate conditions for large-scale commercial deployment.
The distinction is increasingly relevant as more AI companies publish model weights without using a conventional open-source license.
Chinese AI darling Moonshot for example, made the weights of its frontier class Kimi K3 model available earlier this week under a custom "open" license that includes additional commercial conditions rather than the comparatively straightforward terms of Apache 2.0.
For legal, procurement and platform teams, that difference can materially simplify adoption. Apache 2.0 does not eliminate the need to review acceptable-use policies, data provenance, regulatory exposure or downstream safety obligations. But it gives organizations a clearer starting point for building internal systems, shipping commercial products and maintaining modified versions of the model.
A more repeatable model-development pipeline
Inkling-Small also shows how quickly Thinking Machines has turned its first large model release into a repeatable engineering process.
“Whereas I felt like it took a village to release Inkling, Inkling-Small felt much more routine 😆 We just took the pipeline used for Inkling, passed in a smaller model, and voila — new model! Inkling Small benefited quite a bit vs Inkling from some minor improvements, but there’s still so much more left in the tank...”
The comment suggests the company is no longer treating each model as a one-off research project. Instead, it is building a reusable pipeline for pre-training, post-training, reinforcement learning, evaluation and release.
Thinking Machines says Inkling-Small benefited from an improved pre-training data mix, changes to the machine-learning recipe and on-policy distillation using Inkling as a teacher. The team then continued agentic coding reinforcement learning for two weeks.
Mira Murati emphasized the same point in her own post, describing Inkling-Small as comparable to Inkling at one quarter of the size and highlighting that the weights were open and fine-tunable on Tinker immediately.
How enterprises and AI builders should think about Inkling Small
The company is also distributing full BF16 and NVFP4 checkpoints and supporting deployment through SGLang, vLLM, TokenSpeed, Unsloth and Hugging Face tooling.
That combination gives developers several deployment paths: use an API, fine-tune through Tinker, rely on a third-party inference provider, or operate the model on private infrastructure.
Inkling-Small is not a model that most individuals will download and run locally. But for businesses deciding between a very large flagship and a more manageable open-weight system, it presents a compelling compromise: nearly the same measured intelligence, stronger results on several coding and reasoning tasks, lower token pricing, a smaller hardware footprint and a license that permits broad commercial development.
The broader signal may be just as important. Thinking Machines is showing that Inkling was not a one-time release. The company is already compressing its model family, refining its training pipeline and moving toward a cadence in which open-weight multimodal systems can be produced, improved and deployed more routinely.
OpenAI is sharply reducing the prices of two models in its GPT-5.6 frontier series, cutting GPT-5.6 Luna, the smallest and fastest model in the series, by 80% and GPT-5.6 Terra, the mid-tier model, by 20%, while adding a premium Fast mode for its flagship GPT-5.6 Sol model.
OpenAI is successfully undercutting Google's price per intelligence and attempting to sway Anthropic users, who may not mind paying more, with a speed boost.
OpenAI says Luna will now cost $0.20 per million input tokens and $1.20 per million output tokens, for a combined input-plus-output price of $1.40 per million tokens.
Terra will cost $2 per million input tokens and $12 per million output tokens, for a combined price of $14.
Pricing for Sol Standard remains unchanged at $5 per million input tokens and $30 per million output tokens. OpenAI is also adding Sol Fast mode at twice the Standard price: $10 per million input tokens and $60 per million output tokens.
The company says Fast mode delivers up to 2.5 times the throughput without changing the model’s underlying intelligence.
Pricing is shown per one million tokens. Total cost is calculated as input price plus output price. Cached-input pricing is excluded to keep the comparison consistent across providers.
OpenAI moves Luna into the low-cost tier
The most consequential change is the Luna price cut.
When OpenAI introduced the GPT-5.6 series, Luna was priced at $1 per million input tokens and $6 per million output tokens, for a combined total of $7. The new pricing reduces that combined figure to $1.40.
That places Luna below Google’s Gemini 3.5 Flash-Lite, which costs a combined $2.80 per million input and output tokens, and far below Gemini 3.6 Flash at $9. Luna also now costs less than OpenAI’s own GPT-5.4 and Terra models by a wide margin.
It is not the cheapest model in the broader market. Xiaomi’s MiMo-V2.5 Flash, DeepSeek’s flash model and several other APIs remain less expensive on a pure token basis. But the reduction brings an OpenAI frontier-series model into direct competition with the market’s low-cost inference tier.
OpenAI says the GPT-5.6 series represents its frontier model family, with Sol positioned at the top of the lineup, Terra as the middle tier and Luna as the smallest and fastest option.
The lineup was initially released in late June 2026 through a limited rollout by U.S. government request, before broader access, with each model intended to offer a different tradeoff among intelligence, latency and cost.
Sol is aimed at the most complex reasoning-heavy and agentic workloads, including advanced coding, multi-step planning and tool-using systems, while Terra is designed for general production use where a balance of capability and efficiency is required. Luna is positioned for high-throughput, low-latency tasks such as summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint.
Terra drops to match Google’s Gemini 3.1 Pro pricing
Terra’s 20% reduction moves its combined price from $17.50 to $14 per million tokens.
At that level, Terra now matches Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less.
It also undercuts OpenAI’s GPT-5.4, which remains priced at $2.50 per million input tokens and $15 per million output tokens, offering the same intelligence for about 1/13th the cost, as Krea AI's Nic Dunz noted on X:
The adjustment creates a wider separation between OpenAI’s three GPT-5.6 tiers. Luna costs one-tenth as much as Terra on a simple combined input-plus-output basis, while Terra costs 60% less than Sol Standard.
Sol Fast moves in the opposite direction. At a combined $70 per million tokens, it is the most expensive model configuration in the comparison below, reflecting OpenAI’s decision to charge a premium for latency-sensitive workloads rather than lower Sol’s base price.
Cuts follow Google’s low-cost Gemini releases and Anthropic's Claude Opus 5
Google priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Gemini 3.5 Flash-Lite costs $0.30 per million input tokens and $2.50 per million output tokens.
Google framed both models around the economics of agent deployment, arguing that lower token usage, fewer reasoning steps and reduced tool calls could lower the total cost of long-running software engineering and knowledge-work tasks.
Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. Gemini 3.5 Flash-Lite is positioned as the fastest model in Google’s 3.5 series.
However, OpenAI's models are more performant than Google's, according to third party analysis outfits like Artificial Analysis, with even the Luna model outperforming Gemini 3.6 Flash and the older Gemini 3.1 Pro model, making the cost-per intelligence much more favorable to OpenAI.
As AI coding startup Cognition noted on X, GPT-5.6 now "sits on the pareto curve of price/performance efficiency," posting an animation of the GPT-5.6 series moving left on a chart representing intelligence on the y axis and cost on the x, showing that the models now offer among the most superior intelligence for lowest cost on the market.
And yet, rival Anthropic's Claude Opus 5 remains about as performant as GPT-5.6 Sol, yet is 6% cheaper.
The model costs $5 per million input tokens and $25 per million output tokens—the same rates as Opus 4.8—but Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 model at roughly half the cost.
Unlike OpenAI’s Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, it effectively lowered the price per unit of capability by replacing Opus 4.8 with a more capable model at the same $30 combined input-and-output rate. Anthropic also added an adjustable effort setting that allows developers to trade reasoning depth for speed and token savings.
That distinction matters for enterprise buyers. OpenAI is directly cutting per-token rates, Google is pairing lower prices with reductions in token use and tool calls, and Anthropic is emphasizing stronger task performance at an unchanged price. All three approaches target the same operational metric: the total cost of completing production work, rather than the advertised cost of an individual token alone.
The timing highlights how quickly pricing has become a competitive lever among frontier model providers. OpenAI’s response does not introduce a new model generation. Instead, it changes the economics of deploying models that were released only recently.
The market shifts from model access to model economics
The cuts indicate that access to frontier-level capability is no longer the only point of competition. The next question for enterprises is how cheaply and predictably those models can run in production.
OpenAI is still not the lowest-priced provider on a pure token basis. But Luna’s 80% reduction materially changes its position, moving it from the middle of the market into a pricing tier populated by smaller models from Google, Xiaomi, DeepSeek, MiniMax and other vendors.
That matters most for high-volume applications, where relatively small differences in token pricing can compound across coding agents, document systems, internal search tools and automated workflows.
OpenAI’s latest move therefore looks less like a routine adjustment and more like a repositioning of the GPT-5.6 series. Sol remains the premium option, Terra moves closer to competing pro-tier systems, and Luna becomes the company’s direct answer to the industry’s growing low-cost model segment.
Every time a Mastercard gets tapped, the network has less than a tenth of a second to judge how likely the purchase is to be fraudulent. It made that call across 175 billion transactions last year. Now the buyer on the other side of that judgment is starting to change, and Greg Ulrich, the company's chief AI and data officer, spelled out the consequence for the VB Transform 2026 audience in Menlo Park on July 14. "We've built a bunch of risk rules over time that were intended to stop a bot from transacting," Ulrich said. "Now we need to enable the bot to transact, so that requires a change to our risk framework and our risk rules."
Ulrich joined Mastercard eleven years ago when an analytics company he worked at was acquired, and said trust struck him from day one on the job. "It's what enables a merchant that's never met you to accept payment and ensure that they're going to get paid. It's what enables you as a consumer to transact and ensure that things are going to work out in a trusted, secure way. And if something goes wrong, there's a safe and secure path for a dispute and to resolve this," he said.
175 billion transactions, scored in under 100 milliseconds
He took the audience inside each of those calls. "When you tap your Mastercard to pay for a product or service, we're providing a score to that transaction," he said. "We have under 100 milliseconds to look at that and give a score from zero to 999 about how likely is that to be fraudulent or real. And we pass that on to the issuing bank."
Generative AI widened what that score can see. "Because we have new technology, we can bring in more data, we can bring in more context, and now we're finding that we can identify 300, 400% more fraudulent transactions at those high-risk bands," Ulrich said, without adding friction or false positives for consumers. The company's Safety Net system has stopped more than 70 billion fraudulent transactions, he told the audience, and Mastercard is building its own transformer model on its transaction data as a foundation for new safety, security, and personalization solutions. VentureBeat's Beyond the Pilot podcast took that production fraud stack apart in detail earlier this year.
A third of the services business already runs on AI
The business stakes reach past fraud. About 40% of Mastercard's company is now based on services, Ulrich said, including marketing services; fraud, safety and security; and business intelligence. "A third of those are predicated on AI, and those are growing at a much faster clip than everything else," he said.
One line he returned to all session went further. "What's going to enable AI to continue to scale is not the capabilities of the agents, it's how much we trust those agents to do on our behalf as a consumer, as a business, as a financial institution, or otherwise," he said.
Five layers stand between agents and the network
Agentic commerce changes the object being secured. "Instead of a single atomic transaction where I say go buy something, I'm effectively delegating authority, or a consumer's delegating authority, a business is delegating authority," Ulrich said. "And when that happens, it's a much more complicated transaction." Trust, in turn, has a precondition. "The only way it's going to work with trust is if we can identify what was the intent, what are the behaviors, what are the constraints that were intended in that transaction."
Ulrich walked through five layers Mastercard has built against that problem. Identity comes first. "I want to make sure I can understand not just who the consumer is, but who the agent is, that I combine them together and that I have KYA or know your agent, that I'm validating that it's legitimate technology, that it's a legitimate agent," he said. "We can register it into our system."
Verifiable intent settles the "wrong-Nikes" problem
Verifiable intent is second, a tamper-proof cryptographic record of the original instructions that travels with the transaction. "If you've asked for Nike black Nikes in size 12, but you got them on a final sale and they're not returnable and that wasn't in your instruction, there's a way to look at that in an objective and clear way on the back end," he explained.
Controls form the third layer, defining which merchants an agent can buy from, at what limit, and under what constraints. Execution runs through Mastercard Agent Pay, which carries "the tokenization, authentication, the acceptance framework embedded within it" and has launched with Microsoft, OpenAI, Google, and others, Ulrich said. Intelligence is the fifth layer, spanning risk rules, insight tokens that grant "consented or permissioned access to insights" for personalized recommendations, and monitoring through Recorded Future to identify threat actors in the system.
The bigger prize is a procurement agent with a budget
Consumer purchases are where agentic commerce started. Ulrich pointed the room past them, to business-to-business procurement as the larger opportunity. His example was a manufacturer that wants an always-on assembly line, with an agent that manages inventory levels, tracks when stock runs low, replenishes automatically, and understands the budget and the approved suppliers. "When you can start enabling that, you require those same five layers for that type of transaction," he said.
Making it work across companies multiplies the parties that have to trust each other. "You need clear standards for identity, you need clear standards for intent, you need these to work across. You're gonna have a procurement agent, a supplier agent, a banking agent. They're all gonna need to communicate to enable this to happen in an autonomous way, and that's gonna require really scaled trust infrastructure."
Powerful new models, same security motion
Mastercard sat in the early wave of Project Glasswing with Anthropic's Mythos model, and worked with OpenAI's GPT-5.5-Cyber, he said. "What we've seen from both of those is incredibly powerful models finding new vulnerabilities in the ecosystem that were difficult to detect previously, but it's really a new tool as opposed to a new motion," Ulrich said.
Inside the company, the chief security officer leads that work. A dedicated team has prioritized the most critical assets, runs them through the models routinely, tracks findings by high, medium, and low severity, and uses the same technology to handle patches. Ulrich said the approach has already been extended out, and that Mastercard is working to make the same architecture and patching available to others as well.
What Mastercard would build differently after 14 months
"The guardrails, the security, all this stuff has to be embedded at the front end. These can't be things that we're adding on at the back end. That's lesson one. Lesson two is you have to be operating for scale, and the other one is around observability and accountability matter as much as the intelligence," Ulrich said, counting off what building inside Mastercard taught the team. The company built what he described as an agentic factory, an operating system with the compliance, the observability, and the guardrails built in rather than bolted on per agent. Model drift, once tracked manually by dedicated teams, is now automated into that factory.
Asked by an audience member about the gotchas, Ulrich did not soften the pilot-to-production trap. "If you're trying to extend that and then add guardrails in as you're extending it, once you've already built it, I think you're doomed to fail," he said.
Mastercard built a series of agents last year for its 4,000 consultants, covering deep research, text to SQL, Excel, and PowerPoint, tools that by his account did not exist at the level Mastercard needed. Were the company starting today, Ulrich said, it would build them fundamentally differently. "I don't know that we anticipated when we built things fourteen months ago that we would be rethinking the fundamental architecture and the approach already."
Agentic identity joins KYB and KYC
The identity layer is where Ulrich expects the market to move next. Inside Agent Pay, Mastercard authenticates the consumer the way it does in traditional e-commerce and binds the agent to that person. "Outside of that framework, I think there will be open standards to identify who an agent is and bind the agent with the consumer," he said. "And then we can tie that with verifiable intent."
He called identity "one of the faster-growing ecosystems," noting Mastercard has been expanding there organically and inorganically for about six or seven years, with the work now spanning "agentic identity as well as the traditional KYB and KYC identity." The risk rules that keep bots off the network came out of more than two decades of applying AI to those transactions. The rewrite, for the agents Mastercard now wants to let in, is already underway on the same network that scored 175 billion of them last year.
Less than a year after emerging from stealth to tackle non-human identity security, Israeli cybersecurity startup Hush Security believes the enterprise AI security conversation has fundamentally changed.
The company, which earlier this week announced a $30 million Series A round led by returning investors Battery Ventures and YL Ventures with Akamai Technologies joining as a strategic investor, argues that organizations are rapidly moving beyond experimenting with generative AI assistants and into deploying autonomous software agents that require an entirely different security model.
While the funding will help expand engineering, U.S. sales and enterprise integrations, Hush is framing the announcement primarily as evidence that identity—not models—is becoming the critical control plane for enterprise AI.
"The discussion has moved incredibly fast," CEO and co-founder Micha Rave told VentureBeat in a video call interview following the funding news.
When Hush launched last year, the company's focus was securing non-human identities—API keys, service accounts, machine credentials and other identities used by software rather than people.
Since then, Rave says, customers have increasingly asked a different question: how do they safely allow AI agents to operate inside production systems? This is a pertinent and urgent question ever since Hugging Face revealed in mid-July it was hacked by an autonomous AI agent, later identified as an OpenAI test agent running internally that escaped its secure sandbox, powered in part by an unreleased model.
The company's original thesis was that enterprises had accumulated thousands of long-lived machine credentials that were difficult to rotate, audit and secure. Rather than relying on static secrets, Hush developed an identity-based system that brokers short-lived, policy-driven access for machines.
Rave says AI agents amplify that same problem.
"Software now acts autonomously, on its own initiative, inside your most sensitive systems," he said. "AI agents need strict identity, not just API keys."
Unlike traditional automation, AI agents frequently act across multiple enterprise systems, invoke external services, make decisions independently and often execute actions using the permissions of the human who launched them. In practice, organizations often grant an agent broad OAuth permissions or administrator credentials simply to enable it to complete tasks.
That creates what Hush describes as an identity problem rather than simply an AI problem.
During the interview, Rave said virtually every security leader he speaks with faces the same dilemma: either slow AI adoption until appropriate controls exist or allow employees to connect new agents directly into corporate systems despite limited governance.
"The answer," he said, "is that they let everything in. You cannot stop innovation in the name of security."
Identity becomes the control point
Rather than treating AI agents as another application requiring credentials, Hush is extending its existing non-human identity platform into what it calls an "Identity Gateway" for AI agents.
The platform sits between agents and enterprise resources, allowing organizations to discover agents, assign each one its own identity, associate it with a responsible human owner, broker task-specific permissions at runtime and maintain centralized audit logs.
Instead of allowing an agent to inherit all of a user's privileges indefinitely, Hush attempts to enforce what it calls "least agency"—granting only the permissions necessary for the specific task being executed.
The company says every action can be logged, attributed and revoked from a single control plane, while administrators retain the ability to terminate an agent's access immediately if necessary.
This represents a broader shift in enterprise identity management. Human identities have long been governed through identity providers, single sign-on and privileged access management systems. Machine identities have increasingly received similar attention as organizations modernized cloud infrastructure. Hush argues autonomous AI agents now represent a third identity category requiring dedicated governance.
Hush has not publicly posted its pricing for the Identity Gateway solution, nor its offerings more generally. But the company did release a Free plan that gives organizations access to runtime visibility for AI agents and non-human identities, risk analysis, and identity-based access controls intended to replace long-lived credentials, with no credit card or time limit required.
Governing every kind of enterprise agent
Hush says enterprises are no longer dealing with a single category of AI software.
During the interview, Rave described three broad classes emerging inside organizations:
Desktop coding assistants and productivity agents such as Claude, Cursor and VS Code integrations.
Enterprise AI platform agents running on services such as Microsoft Foundry, Salesforce Agentforce or AWS AgentCore.
Custom agents organizations build internally for business processes or customer-facing applications.
Each introduces different governance challenges, but all ultimately require controlled access to enterprise systems.
The problem, according to Hush, is that many agents currently authenticate using inherited human credentials or long-lived API keys, making it difficult to determine whether an action originated from a person or from an autonomous system acting on that person's behalf.
"If I see something in the Salesforce logs," Rave said during the interview, "did the user do that, or was it the agent the user was using?"
That attribution challenge becomes increasingly significant as organizations begin deploying multiple autonomous systems capable of initiating actions without direct human approval.
Existing identity tools weren't designed for AI agents
Rather than replacing identity providers or secrets managers, Hush positions itself as filling a gap between them.
Traditional IAM platforms authenticate employees. Secrets managers store credentials. Neither, the company argues, governs the runtime behavior of autonomous software acting on behalf of humans across multiple systems.
Hush says its platform continuously discovers known and shadow agents across enterprise environments, assigns ownership, brokers just-in-time credentials and records every interaction in a centralized audit trail. According to its product documentation, organizations do not need to modify their existing agents because the platform operates by brokering access requests rather than changing application logic.
That identity-first approach is attracting customers already deploying enterprise AI initiatives.
IT infrastructure services provider Kyndryl says it has deployed Hush internally and has begun offering the platform to enterprise customers.
"Our collaboration with Hush is rooted in a shared security philosophy: identity is the ultimate control point for the modern agentic workforce," said Adeel Saeed, senior vice president and CTO for Global Cyber Resiliency at Kyndryl, in a prepared statement.
Akamai's participation in the funding round similarly reflects what the company sees as an architectural rather than incremental shift.
"AI agents are driving the next transformation, and identity is the piece most companies haven't solved yet," said Ramanath Iyer, Akamai's chief strategist.
Security priorities are moving beyond the model itself
The broader AI security market has spent the past two years focused largely on prompt injection, model vulnerabilities, jailbreaks and LLM safety. Those remain active research areas, but enterprise deployments increasingly face operational questions around what autonomous systems are permitted to access and how those actions can be governed.
Hush argues that identity is becoming the enforcement layer for answering those questions.
Rather than asking whether an AI model can safely generate code or summarize documents, enterprises increasingly need to determine which systems an agent may access, whose authority it exercises, how permissions are delegated, and how every action can be traced back to an accountable owner.
Whether Hush's identity-centric approach becomes the dominant model remains to be seen. But as enterprises move from experimenting with AI assistants to deploying thousands of autonomous software agents, the company is betting that the next major security challenge won't be securing the models themselves—it will be securely managing the identities of the software acting on their behalf.
VentureBeat’s June research found that 69% of enterprises are still running AI agents that share credentials, a practice associated with higher rates of security incidents and near-incidents.
But at VB Transform 2026, Mukesh Karki, CTO of NTT DATA AIVista, and Mayank Upadhyay, chief security and trust officer at Snowflake, argued that fixing identity is only the first step. Enterprises also need action-level authorization and tamper-resistant audit trails built into every agent interaction if they’re going to deploy autonomous systems safely at scale.
"These organizations need to be able to prove to their auditors in a very tamper-resistant fashion that those records showing what they did actually prove what they're doing," Karki said. "And the provability is essentially your license to operate in a regulatory environment."
Why shared credentials cause agentic AI security incidents
The problem, Upadhyay says, is many assumptions were carried over from an earlier generation of software.
"In the traditional software world, a human being clicks somewhere and the software does something very deterministic, and you know which API it's going to call," he said. "But in the agentic world, the software has a brain of its own, and it's constantly rewiring itself. If you give this software more permission than it needs for a particular goal, agents are exploratory by nature, so they're going to try lots of different things, and you'll have unintended side effects."
Embedding a single static API key compounds the exposure, he added.
"It's a really bad pattern if you have one API key, you shove it into the agent, and it's talking as anybody to a particular SaaS service, because then you're giving this agent the union of everybody's needs," he said, noting that the second failure mode is forensic, since "things may go wrong, and you wouldn't be able to attribute it to the right agent."
Scoped credentials are only the starting point in regulated industries
Karki, whose clients are mostly in insurance, healthcare, and finance, treats scoped credentials as table stakes.
"In a regulatory setting, an agent that's not broadly scoped with shared scope credentials is not going to run, period," Karki said. "Having a scope credential is just a starting point. There are actually two layered constraints. One is the jurisdiction in which the agent operates, and then it's the jurisdiction or the rules of that organization."
For instance, a claims adjustment agent in Washington State operates on different regulations than one in California, he adds, and every claim is different.
"Those scoped credentials are not enough, because it has to be action-based and rules-based at the time it's taking action," he added.
Where the employee analogy for AI agents breaks down
The employee analogy, Karki argued, only goes so far. Agents still need to learn an organization’s unique context, much as a new employee does. But unlike people, enterprises can’t realistically build trust with thousands of agents over time.
“A star employee in one organization might not be the best employee when they move to a different organization, not because they became worse, but because they don’t have the context of this new place, and the same is true with agents,” Karki said. “If every employee has 100 agents, you can’t say you’re going to onboard these agents and do a background check on them.”
Upadhyay said the employee analogy should place agents one rung lower in the organizational hierarchy.
"Treat them like interns," he suggested. "They have good intent, but they don't always know what they're doing, and you have to keep your eye on them while you gradually build trust."
On the Snowflake platform, administrators can impose platform-wide guardrails such as read-only operations, while developers further narrow an agent’s permissions when they launch each session.
A three-layer approach to AI agent governance
There's no question where governance belongs, Karki says.
"Governance has to happen at every agent action, and it has to sit outside the agent," he explained. "That's the only way you'll be able to prove later that the agent took an action it was allowed to take."
Upadhyay broke governance into three layers:
The agent layer covers identity, tool permissions, and MCP governance.
The model layer addresses indirect prompt injection and enables models to run inside the customer’s VPC so prompts remain invisible to the model provider.
The data layer covers least-privilege access, zero-copy architecture, and role-based access control.
For agents to work properly, governance is required across all three.
What enterprises should audit first
For enterprises auditing the governance of existing AI agents, Upadhyay recommends starting in two places. The first is auditing permissions for static secrets, the largest fixable attack vector. Next is addressing shadow AI through an MCP gateway, so developers no longer have to run bootlegged open-source MCP servers under their desks and administrators have visibility into who’s talking to which MCP server.
There's a tradeoff between constraint and capability, and that can be addressed at the task level, with confidence scoring used to withhold autonomous execution on high-risk actions, and sandboxing as a middle path. But Karki cautions enterprises already scaling their agentic systems.
"A lot of this can't be retrofitted after you have an agentic system running, and it's even harder to retrofit if you have to prove to your auditors why exactly the agent behaved the way it did," he explained. "Provability has to be built ground up when you're designing the system."
Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
Enterprise AI has moved from experiment to execution, and that shift is beginning to show real returns. The SAP Value of AI Report 2026, produced with Oxford Economics and based on a survey of 2,600 business leaders across 13 countries, found that AI now supports nearly one-third of all tasks in the average organization, rising to 30% from 25% last year.
ROI expectations for agentic AI have jumped from 10% last year to 17% this year, but many organizations believe AI could be delivering far more value. The report reveals that the gap comes down to strategy, data, and governance, rather than access to the newest model, says Sean Kask, chief AI strategy officer at SAP.
"AI has moved from experiment to execution, and that's beginning to show real returns, but there's still a long way to go," Kask says. "That's because AI that lacks context, whether that's processes, data, or governance, at best creates activity without outcomes and at worst creates risk."
Companies are still taking a piecemeal approach to AI
Even as investment accelerates, more than half of organizations still invest in AI in an ad hoc or piecemeal way, and only 17% report a strategic, holistic approach to prioritization, though that figure has nearly doubled from 9% a year ago.
That fragmentation may go back to board-level demands that employees start adopting AI without a strategy or adequate AI literacy behind it, which could produce scattered skunkworks efforts. In other companies, a lack of attention at board level can leave employees bringing their own tools to work and just experimenting.
"You end up with a lot of organic, disjointed AI initiatives that pop up, and they struggled sometimes just because of data quality," Kask said. "But even the initiatives taking a strategic approach are still working in silos, where they may have consistent data that works in that one use case, but they're still not at the level where they're transforming an entire business process."
That may help explain one of the report’s more counterintuitive findings: 69% of businesses say they are satisfied with their AI ROI, because they've proven AI can generate returns. Yet 67% remain unconvinced the technology is delivering its full potential, because that learning experience has made them aware of both how much more value AI can deliver and the challenges they need to overcome to scale it.
Agents are changing the economics of enterprise AI
SAP shipped more than 400 AI use cases across its portfolio so far, with many more in the works. Agents represent the next expansion, because they can plan and reason through multiple steps and tools to reach an objective, which mirrors how people and processes work, Kask says.
"You're giving a task or an objective to an AI system, and it's able to iteratively work through several steps and access various tools to achieve that outcome," Kask said. "For instance, we've released, in beta, an agent for accruals accounting, a job that would typically take an accountant around 12 hours a month for a mid-size-company, and it gets reduced to two or three hours. So now scale that out across all these processes and its huge potential."
In fact, general AI ROI went from 16% to 21% this year, and should grow to $15.9m in two years’ time, even as only 3% say they are fully prepared for it.
Data quality remains the biggest barrier to AI value
Getting ready for agents comes down to two fundamental requirements: connecting agents to contextually rich data, and governing them at scale. Data quality and availability are now the number-one reason organizations say they're not getting more value from AI, according to 73% of respondents, with 79% reporting rework, delays, or backlogs from low-quality outputs at least occasionally.
The nature of the problem has changed compared to classic deep learning. Foundation models eliminate much of the need to find data, extract it, clean it, and train bespoke models, but they make preserving business context far more important.
"As soon as you extract data from an ERP system, you break all the contextual information, all of the semantics, and for generative AI, that's the most useful part," Kask said.
SAP is able to preserve that context at scale through a knowledge graph in its cloud ERP that maps 452,000 ABAP tables and 7.3 million data fields. In SAP Business Data Cloud, data products present information such as invoices and suppliers consistently across SAP and non-SAP systems without losing their business meaning.
AI governance is the biggest challenge companies don't know they have
As AI becomes more deeply embedded in business processes, governance is emerging as the next enterprise challenge. Only 12% of businesses say they are fully prepared to govern AI, while 69% acknowledge occasional to frequent use of unapproved shadow AI tools.
“As companies roll out their AI initiatives, they often discover shadow agents – agents that can access data they shouldn’t or take actions they shouldn’t. The question then becomes: How do we audit these things?” Kask said.
SAP’s AI Agent Hub responds by discovering and creating an inventory of agents, LLMs, and MCP servers, and customers have already surfaced thousands of SAP and non-SAP agents inside their landscapes that they did not know they had. It then layers on lifecycle management, identity and access control, and performance monitoring. Kask compares the discipline to hiring, since most companies would never onboard an employee without knowing which access rights and permissions that person needs to have in their role.
Governance, however, extends beyond technology. Workforce transformation runs alongside the data work, with almost 80% of respondents agreeing that maximizing AI value requires more than technical upskilling and 75% already planning to reskill employees. The conversation is shifting away from which jobs AI will replace and toward how people and AI collaborate most effectively, since agents still require human oversight, redesigned workflows, and stronger judgment.
All of this points toward what SAP calls the Autonomous Enterprise, which connects agents to contextually rich data and enterprise governance across functional silos while using Joule as the natural-language, generative interface between people and systems.
“Realizing real value from AI is not going to be easy because it demands a new approach,” Kask concluded. “It is ultimately a human change more than a technical one, because you can only achieve real value if agents, processes, and people work as one.”
Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
A security team approving an open-source model for production today starts with a repository page. The page lists the model name, the license, and a tag identifying the base model it descended from. That tag is a string the uploader typed. Hugging Face does not require uploaders to substantiate the claim through weight-level analysis.
The ATOM Report, published by Nathan Lambert and Florian Brand at Interconnects AI in April 2026, tracked roughly 1,500 mainline open models. ATOM identifies derivatives through the Hugging Face base_model tag, a field the uploader populates, filtering to models whose base model appears in the tracked list and that have more than five lifetime downloads and excluding GGUF and MLX re-uploads. By that measure, Alibaba’s Qwen family is the declared parent of 69% of new open-model derivatives as of February 2026, up from 1% in January 2024. Chinese labs overall account for 70%. Europe sits at 4%. Cumulative tracked downloads across the three regions reached 2.04 billion through March 2026.
The verification gap extends to scan coverage. Cisco Foundation AI scans every public file uploaded to Hugging Face through an updated ClamAV engine, and the platform surfaces a file-level badge per file. Hugging Face’s own malware scanning documentation notes a file with neither an ok nor an infected badge may be queued, still scanning, or errored. At a given review point, a repository may contain files without completed scan results. Coverage has been an assumption, not an attribute anyone could read before approving a model.
From command line to public lookup
Cisco on Thursday published the AI Supply Chain Provenance Explorer, a free public database covering almost 900 open models. Each entry can carry provider headquarters, a fingerprinted lineage graph, license restrictions, and a files-scanned count. The tool extends Cisco’s Model Provenance Kit, an open-source Python toolkit released in April that fingerprinted roughly 150 base models across 45+ families and 20+ publishers. Coverage grew roughly sixfold in a quarter.
The April release was a command-line tool. Running it meant a local Python environment, downloading model weights that run into tens of gigabytes, and dedicating engineer hours per model. The Explorer queries results Cisco already computed. On Thursday, verifying parentage starts with a search bar, and cost is why enterprises run open weights in the first place.
Amy Chang, head of AI Threat Intelligence and Security Research at Cisco, has been building the case for why verification gaps matter. During a VB Transform 2026 agentic security panel, Chang presented findings from 6,986 multi-turn attacks against 15 flagship models, with success rates reaching 88.3%. "If you don’t understand how models are susceptible to different types of attacks, then you are unable to account for how that model that is powering your agent, that is powering your application, to understand where those failure points are," Chang told the audience. Understanding failure points starts with knowing which model you are running.
The Explorer also surfaces data Cisco already uses operationally. The company’s Cerberus system inspects models entering Hugging Face and feeds Secure Access policies that block by risky license or region of origin. The Explorer makes that class of information free and searchable without a Cisco product.
How fingerprinting replaces the tag
The Explorer grounds model relationships in similarity scores rather than self-reported metadata. Cisco’s Model Provenance Kit works in two scored stages. Stage one compares architecture metadata before loading any weights. When metadata is ambiguous, stage two extracts five weight-level signals. Embedding Anchor Similarity captures geometric relationships that survive fine-tuning. Embedding Norm Distribution encodes word frequency patterns. Norm Layer Fingerprint reads layers stable across fine-tuning. Layer Energy Profile compares distributions across network depth. Weight-Value Cosine directly compares weight values, and independently trained models show essentially zero correlation on this signal. Cisco reported 96.4% accuracy on its own 111-pair benchmark at a 0.70 threshold, with an F1 of 0.963. Four pairs were misclassified, all involving extreme architectural transformation that Cisco calls a fundamental limit of pairwise weight comparison.
Tokenizer signals are computed for diagnostics but deliberately excluded from the provenance score. StableLM and Pythia both use the GPT-NeoX tokenizer and would score as related despite sharing no weight lineage. Excluding tokenizer data prevents false positives.
Behavioral fingerprinting adds a second approach. Jonah Leshin, Manish Shah, and Ian Timmis at Project VAIL, working with Daniel Kang at UIUC, published work on behavioral endpoint stability showing that a model endpoint can stay healthy while its effective identity changes through weight updates, quantization, or routing. Cisco’s launch blog states the Explorer integrates both static fingerprinting and behavioral-similarity analysis to ground the lineage graph. Static analysis supplies weight-level evidence of training-time derivation. Behavioral analysis catches runtime identity drift.
Where existing tools fall short
The Explorer carries real limits. Almost 900 models is a meaningful start, but Hugging Face hosts more than 2 million as of spring 2026. Models outside the boundary still depend on the self-reported tag. Cisco has not said whether the Explorer exposes an API, and without one, a team can look models up by hand but cannot wire the check into a CI gate. That is the line between a governance artifact and a control.
Traditional SCA tools face a structural mismatch because they were built for dependency manifests and container images. Sakshi Grover, senior research manager for cybersecurity at IDC, said in CSO Online that traditional SCA "was designed to inspect dependency manifests, libraries, and container images" and "is far less effective at identifying" the risks tied to AI workflows. Gartner director analyst Jaishiv Prakash told the same outlet that enterprises need "dedicated controls for model sources, approved versions, access, and runtime validation at the registry layer." Both were commenting on broader supply chain risks, but the gap they describe is the one the Explorer targets.
Cisco’s Model Provenance Constitution defines where one model counts as a derivative of another. The constitution defaults to labeling ambiguous pairs as independent, because a false positive triggers a licensing accusation while a false negative gets caught during manual review. That deliberate conservatism supports the 96.4% accuracy figure. Derivation is not binary, and fingerprinting is one form of evidence alongside documentation and checkpoint verification.
What goes in the approval record
On August 2, the European Commission gains its AI Act enforcement powers over GPAI model providers, with fines up to 15 million euros or 3% of global turnover, whichever is higher. Organizations that substantially modify and place an open model on the EU market can acquire provider status, with Commission guidance treating modification compute exceeding one-third of the original’s. The Act’s open-source exemption under Article 53(2) requires a genuinely free and open-source license permitting access, use, modification, and redistribution, with weights, architecture, and usage information all public. Public weights alone do not qualify. Llama’s community license carries a monthly-active-user threshold and a disqualifier the Commission guidance names explicitly. Llama and Gemma together account for roughly a fifth of new derivatives in the ATOM counts, and both carry licenses the Commission criteria would likely disqualify. License classification becomes part of the provenance review, and that is exactly what the Explorer surfaces.
The board question that arrives first after a base-model vulnerability disclosure is straightforward: "Which of our production models inherits this weakness, and how do we know?" The answer today requires a manual hunt through repository pages, tracing self-reported tags that no weight-level analysis has confirmed. The Explorer converts that hunt into a lookup for the models it covers.
Four fields belong in the approval record that most organizations do not carry today. Fingerprint-supported derivation grounded in weight analysis rather than a self-reported tag. A files-scanned count replacing the assumption of coverage with a measurable scan count. Provider headquarters as a filterable field, recognizing that headquarters alone does not resolve export-control exposure, since ownership and deployment location also govern the screening. And license lineage surfaced so legal teams can identify potential upstream terms before a model reaches production.
Cisco released the Supply Chain Provenance Explorer today, and it is available at provenance.aidefense.cisco.com. The database is free, public, and does not require a Cisco product or account.
What changes for a security team on July 30
What the team has today
What the Explorer publishes
Recommended action
Blast radius after a base-model vulnerability. The model name and the base_model tag. Scoping which models inherit a disclosed weakness is a manual hunt through repository pages.
Lineage grounded in similarity scores using two scored stages of fingerprinting on architecture metadata and five weight-level signals. The kit scored 96.4% accuracy at the 0.70 threshold.
Attach fingerprint-supported derivation to each model in the asset inventory so a disclosure triggers a scoped review instead of a hunt.
Malware scan coverage. A file-level badge per file. At a given review point, a repository may contain files without completed scan results. Coverage has been an assumption.
Files-scanned counts and reported malware or unsafe-file findings per model, from ClamAV-based scanning. Scan coverage becomes readable before approval rather than inferred from a badge.
Replace the assumption that a model was scanned with the recorded count. Where coverage is partial, document whether the gap is acceptable and why.
Provider jurisdiction. An organization name on a repository page. A derivative several steps from its origin displays the uploader, not the ancestor.
Provider headquarters, website, and associated HF organizations as a filterable field. Headquarters alone does not resolve export-control exposure.
Add jurisdiction to the approval record. Any team that substantially modifies and places an open model on the EU market faces potential provider obligations under the EU AI Act.
License obligations. A license tag describing what the uploader believes applies. Terms from a base model upstream may not appear on the page the engineer reads.
Common limitations per model, including attribution, non-commercial terms, geographic restrictions, and prohibited use cases. Fingerprinted lineage helps legal teams identify potential upstream terms.
Route license lineage to legal before production, not after a contract references it. Document the position at approval rather than reconstructing it during a dispute.
Few companies face higher stakes when deploying AI than Waymo, the self-driving car company under Alphabet that spun out of Google. Its models do not merely generate text or automate back-office tasks: They help vehicles navigate unpredictable streets, respond to human drivers and make split-second decisions in the physical world.
But the methods Waymo uses to manage those risks — continuous evaluation, carefully curated data, human oversight and clearly defined business outcomes — offer a broader playbook for enterprises deploying AI agents in nearly any industry.
Manasi Joshi, Waymo’s director of engineering for systems intelligence and machine learning, explained at VB Transform 2026 how the autonomous vehicle company trains, tests and deploys AI at scale. To date, Waymo has driven more than 220 million fully autonomous, or "rider-only," miles, with 17 times fewer serious crash injuries than human drivers over the same distance, according to the company.
To achieve these impressive results, Joshi said Waymo has adopted what she called “eval-forced development” or “eval-centric development,” making evaluation a core part of engineering rather than a final check performed before deployment.
“The stage at which our projects are maturing can be easily kind of transpired based on the eval maturity that they showcase,” Joshi said.
In practice, Waymo assesses a project’s readiness partly by examining the maturity of the tests surrounding it. That approach has clear implications for enterprises building customer service agents, coding assistants, financial systems or other AI applications: If a company cannot reliably measure a system’s performance, it may not be ready to place that system into production.
Evals must continue after launch
Joshi said much of Waymo’s quality work has shifted toward evaluations, including tests conducted during model training, after training and inside open-loop and closed-loop simulations.
“Eval is not a one-time task to launch a model,” she said.
Waymo instead treats evaluation as a continuous process spanning driving, simulation and validation. Its methodology combines datasets, performance metrics and infrastructure capable of operating efficiently at scale.
For enterprises, that means testing an agent before launch is insufficient. Teams must continue evaluating it as underlying models, business processes, user behavior and incoming data change. Those evaluations should also connect to actual business outcomes rather than relying solely on broad industry benchmarks.
Joshi cautioned that model-quality measurements are only as trustworthy as the evaluation data behind them. Waymo therefore pairs its performance claims with information about the properties of the datasets used to test its systems.
Testing the rare and dangerous cases
Waymo’s evaluation hierarchy remains grounded in one overriding objective: safety.
The company draws on first-party driving logs, some third-party data and realistic simulations that expose its systems to scenarios spanning billions of synthetic miles. Task owners choose specialized data and metrics for situations involving vulnerable road users, railroad crossings, construction zones and other complex environments.
The same principle applies outside autonomous driving. Enterprises need to test not only the routine requests their agents handle successfully, but also uncommon situations where errors could create financial, legal, security or reputational damage.
Joshi emphasized that Waymo does not leave release decisions entirely to automated systems. Its production-readiness reviews include extensive human oversight, while internal safety leaders approve software releases and service-area expansions.
“This is not AI-driven and completely automated and zero human oversight,” she said. “Human lives are at stake.”
Efficiency cannot come at the expense of reliability
Waymo faces another problem familiar to enterprise AI teams: Demand for compute, storage, memory and network capacity is growing faster than the resources available.
The company pursues efficiency across data extraction and storage, distributed model training, model distillation, simulation and evaluation. It also emphasizes “data efficiency,” selecting the most useful training examples instead of treating greater volume as inherently better.
Waymo began using transformers in 2017 and subsequently expanded into large language models, vision-language models and vision-language-action models. Joshi said the company now uses generative multimodal models as part of its foundation-model strategy.
Waymo divides its technology between onboard systems inside each vehicle and off-board infrastructure used for model development, data processing and simulation. That combination forces the company to optimize both real-time inference and the larger systems supporting it.
Agents need their own evals
Waymo also uses AI agents internally as productivity tools for engineers. Joshi said agents help analyze data distributions, assess data efficiency and triage problems found in vehicle telemetry, training runs and failed evaluation jobs.
The goal is to accelerate investigative work so engineers can devote more time to judgment and difficult technical problems. But Waymo also evaluates those agents to ensure they produce trustworthy, accurate results rather than sending employees down unproductive paths.
For enterprise leaders, Waymo’s larger lesson is that agentic AI requires more than choosing a powerful model. Organizations need a clearly defined objective, representative evaluation data, continuous testing, infrastructure that can operate efficiently and named human decision-makers who remain accountable for deployment.
"Earning trust is supremely important," Joshi said.
Enterprise AI agents can do the work — but the infrastructure to let them talk to each other, prove they should be trusted, and be audited when something goes wrong is still being built.
Here's a look at how five startups are tackling that gap — around orchestration, observability, connectivity, and security — as shown at VB Transform 2026.
BAND is orchestrating all the agents you have running in the background
In the very near future, agents will be deployed everywhere, and they will do work on our behalf, noted Vlad Luzin, CTO and co-founder of BAND.
As he describes it: They will receive tasks, visit registries, recruit other agents to help them, delegate subtasks to AI peers in a “conversational space,” gather and share results, then return a summary to the human user.
BAND is building a coordination infrastructure layer for multi-agent AI systems to make this a reality.
Why don’t Telegram, Slack, or Discord solve the problem? These platforms were built for humans, Luzin noted. Agents have to be onboarded manually in numerous steps, and they can’t see each other; “they are still alone in a kind of digital solitary confinement.”
Similarly, Claude is stateless, and devs often have multiple sessions open at a time that they toggle between for different tasks — something Luzin said creates real friction.
The challenge is connecting remote processes, which Luzin called a distributed systems problem.
“The transportation layer needs to be solved first, how the agents communicate in real time,” he said. Conversations can’t happen through IPs and URLs; they need to be bumped to the abstraction layer so agents can talk across channels, conversational spaces, and platforms.
“Agents see each other. They understand. They can collaborate together. They discuss issues. They fix issues, and they ask for review from another,” Luzin said.
BAND supports autonomous workflows that can run for eight to 20 hours and is compatible with A2A and MCP protocols, according to Luzin. Importantly, humans can join the conversation as agents converse and discover one another, he said.
“We can record and show you all the tasks that your agent generates in real time,” Luzin said.
Conifers is helping defenders move at machine speed
The biggest challenge defenders face today is that they’re still running at human speed, but adversaries are running at machine speed, said Tom Findling, CEO and co-founder of Conifers.
Attackers are already adopting agents, Findling said, and they only have to be successful once to penetrate an enterprise. Malicious campaigns that used to take months and weeks now take hours, even minutes. Security operations, on the other hand, are fragmented, manual, inefficient, and slow.
Findling said Conifers has taken various components of cyber defense — private intelligence, hunting, detection, engineering, investigation, response — and made them agentic. They then broke down the silos between them, he said. Various agentic systems can communicate with one another to ensure that operational defense and active defense are always on and adapting.
Findling said that Conifers’ system is condensing containment time from 7 hours to 12 minutes, and that the company can turn around complex cyber investigations in four minutes or less.
He emphasized the importance of connecting to an enterprise’s existing security tools, whether that be endpoint detection and response (EDR), security information and event management (SIEM), posture management, or others. Conifers helps customers understand their security posture, pain points, which controls are working and which are not, and the areas to invest for the best ROI.
“The threat landscape is changing, detection stays the same, and threat intelligence is not being operationalized,” Findling said. “This is a job for agents.”
Raindrop AI creates an agent audit log
One of the defining problems of the current era is finding critical issues in AI agents, says Ben Hylak, CTO of Raindrop AI.
It’s what he called a “double whammy”: As agents become more capable, complexity increases, as do timelines; they are running for hours or days in some cases. Secondly, issues become catastrophic in sectors like healthcare or defense.
“This problem is getting a lot worse as models and agents improve,” Hylak said, “and I think there's good reason to believe it will continue to get worse.”
Raindrop AI's platform finds critical issues in agents in production and simulates fixes based on past user behavior, Hylak said. That lets teams confirm a fix works as intended before it's live, without introducing unexpected side effects.
The startup’s reinforcement learning (RL) platform optimizes harnesses and trains models directly from Raindrop data, he said. Its pre-deployment simulation engine helps identify what fixes would actually impact in production; its live A/B testing then shows those changes in action.
Messages, tool calls, retries, and errors are captured in one place, and human users are notified (typically via Slack) when there's an issue, he said. Models are trained for every customer, and signals are powering continual learning across models and harnesses. “It is condensed into something that is actually navigable, easy to understand, easy to verify,” Hylak said.
Arcade gives agents the security clearance they need to take action
AI agents are designed to do all kinds of things for you, but they often hit three major snags: authorization, governance, and reliability.
To act on behalf of real users with real permissions, agents need a new type of security architecture, said Sam Partee, co-founder and CTO of Arcade.dev.
Partee said his company’s secure agent runtime provides this authentication and authorization layer so agents can pass critical security reviews. It also provides observability so human users can watch everything an agent is doing. Actions are attributable to the exact moment in time with the least amount of privileged scopes.
Arcade is available in an installable plugin that can be deployed on-prem in a clean room-like environment; companies can continue to use their own sign-in and security tools, Partee said. Whenever anything is run in Arcade, it's gated by the same role-based access controls (RBACs), intrusion detection and prevention systems (IDPS), policies, entitlements, and other already-established checkpoints.
Arcade is tackling the supply chain attack problem, which has “gotten so rampant; it's unbelievable,” Partee noted. Security and observability have continued to be challenging because “largely, the abstraction has been wrong.”
Omilia is tackling the "not straightforward" CX problem
Solving enterprise customer experience (CX) is “really not straightforward,” said Claudio Rodrigues, CPO of Omilia.
Heuristic-based systems are controlled but slow; agentic systems are fast but unpredictable, Rodrigues said. Omilia built its platform to deliver both control and speed together.
The agentic, self-learning offering is built on a philosophy of observing customer service operations as they actually happen, rather than in the abstract. Omilia's agents observe problems first-hand, listen to every customer and agent interaction, ingest data, API specs, screen recordings, and standard operating procedures (SOP), then map those to use cases for customer support, he said.
Contact centers should be a revenue driver, Rodrigues said, and Omilia’s differentiator is its speech-to-text systems and governance and observability layers.
AI creates insights, suggests improvements, automatically generates conversational agents, pulls information from documents and APIs, and designs dialogue flows. Human experts can then test real and simulated interactions and deploy into production under their supervision. Omilia combines all of this into one enterprise-wide engine that continuously learns over time, Rodrigues said.
Rodrigues said the company handles more than 3 billion calls a year, 1 million-plus voice calls a day in some deployments, and has seen 30 to 45% improvement in time to resolution (TTR). Omilia’s agents generate 21x more upsell revenue versus human agents, he said.
In a mature deployment, automation “easily” reaches 80 to 90%, he said. However, “human in the loop is still very fundamental for us.”
Nimble, a New York City-based tech startup VentureBeat previously covered for its efforts to re-invent web search for enterprises by using multiple AI agents to improve accuracy and depth, is taking another step toward its vision of a world in which agents do most of the web searching instead of us typing and reviewing the results manually.
Nimble today launched Web Search Agents, a new retrieval system designed to help AI agents perform more 21% more accurate web research while using significantly fewer tokens — 51% less compared with leading AI search alternatives on comparable, according to the firm.
While Nimble did not disclose its specific benchmarking methodology or competitors evaluated, the results underscore a growing trend in enterprise AI: optimizing retrieval has become as important as improving the underlying language models themselves.
Nimble's leadership says the product combines self-learning retrieval strategies, proprietary web indexes, and live web access to deliver domain-specific search capabilities that outperform general-purpose web search services for enterprise workloads.
"Our research team built self-learning retrieval algorithms that learn a customer's domain," said Nimble CEO and co-founder Uri Knorovich in an interview with VentureBeat. "They find the exact information more efficiently, reduce the amount of multi-hop reasoning required, and lower token usage while improving accuracy."
Rather than positioning itself as another general search engine, Nimble is targeting developers building autonomous agents that require continuously updated information from the public web for research, lead generation, competitive intelligence, compliance, and other business-critical workflows.
It's also designed to slot in seamlessly to an enterprise's existing systems and workflows.
"You can run the agent directly through the Nimble API with zero infrastructure," Knorovich said. "For large enterprises, we're partnering with Microsoft, Oracle, Snowflake, and others so customers can deploy these agent systems inside their own infrastructure."
How does it work and stack up to other, existing AI-powered search and agentic systems? Read on to find out.
Moving beyond generic AI web search into specialized search agents that fit your enterprise's needs
Most AI applications today rely on general-purpose search application programming interfaces (APIs) for search engines and public knowledge bases that return broad collections of files, leaving the language model responsible for determining which sources are relevant.
That process often requires multiple retrieval steps, additional reasoning, and significant token expenditure before an agent produces an answer. This is obviously inefficient and raises the cost spent to run AI search looking through irrelevant sources.
Nimble argues that before long, every enterprise will need its own methods for searching, retrieving, and validating external information since each enterprise relies on its own distinct preferred sources, signals, and standards of trust.
As such, instead of applying one search strategy to every workload, Nimble's Web Search Agents are designed to learn the characteristics of a specific domain and adapt how information is retrieved, providing agents with structured, relevant context rather than forcing them to sift through large amounts of generic search results.
"Instead of one generic retrieval model, we build specialized retrieval models for each customer's domain, making them faster, cheaper, and more accurate," Knorovich explained. "A single enterprise can run hundreds of different agents. Each one has its own domain expertise, guardrails, goals, and search algorithm. The optimization starts with the second search, without requiring any setup from the customer."
Its goal is not only to reduce redundant retrieval, but also to shorten multi-step research paths and avoid repeatedly sending raw pages through a language model for parsing, resulting in the 51% reduced token figure the company cites.
The distinction is particularly relevant for long-running enterprise agents performing research over hours or days rather than answering simple consumer questions. In those scenarios, reducing unnecessary tool calls can significantly lower operating costs while improving answer consistency.
That emphasis reflects a broader shift occurring across the AI tooling ecosystem. As foundation models become increasingly capable, infrastructure vendors are competing on everything surrounding the model—including retrieval, orchestration, memory, observability, and governance.
The company’s latest release extends that vision with a concept it calls “Harness as a Tool,” which powers its new domain-specialized Web Search Agents. Rather than requiring engineering teams to assemble separate search APIs, browser automation, extraction pipelines, validation logic, memory systems, and orchestration code, Nimble packages those capabilities behind a managed interface.
The harness can determine what to search, navigate pages when conventional indexes are insufficient, extract relevant information, validate the results, and return the final context in a form designed for downstream agents.
Nimble also says the system retains domain-specific memory and builds proprietary indexes that improve as customers run more searches.
"The biggest research breakthrough is adding semantic memory and a caching layer to the agent," Knorovich told VentureBeat. "The agent learns usage patterns and domain expertise over time, so every subsequent search becomes faster and more efficient."
As for what domains Nimble can tackle, the company says it can address virtually any knowledge work domain.
"We've seen customers build investment banking analysts, competitive intelligence agents for product managers, go-to-market research agents, newsroom monitoring, insurance applications, life sciences research, and supply chain optimization," Knorovich said. "Our customers surprise us every day with new agent use cases."
However, for enterprises concerned about data privacy and retention, Knorovich assured VentureBeat that: "Nimble is zero-data-retention by design. Customer queries are never stored in our environment, and when customers deploy semantic memory and self-learning models, that knowledge stays in their own tenant—not ours."
Customer deployments point to operational gains
Nimble supported the announcement with early customer examples from AI-native software vendors and enterprise users.
AI-native CRM company Rox reported achieving a 20× reduction in token costs after adopting Nimble’s retrieval infrastructure while simultaneously improving the quality and completeness of information available to its AI agents.
Although the company did not disclose detailed workload measurements or a reproducible baseline, the example illustrates the operational savings retrieval optimization can provide for high-volume agent deployments.
Nimble says its infrastructure currently supports more than 90 million searches each day across Fortune 500 enterprises and AI-native companies operating mission-critical workflows where accuracy, completeness, and enterprise control are essential.
API, SDK and MCP support target AI builders
The platform is immediately available through an API, SDK, and Model Context Protocol (MCP) integration, allowing developers to connect Nimble directly into AI agents regardless of the orchestration framework they use.
Developers can use the platform for several categories of web intelligence, including:
Low-latency live web search
Deep multi-step web research
Web crawling
Structured dataset generation
Domain-specific information retrieval
The company also provides documentation and pre-built agents for common web extraction tasks while allowing developers to build custom retrieval agents using natural-language descriptions instead of manually maintaining scraping logic.
Nimble is offering two notably different consumption models. Developers can begin with a pay-as-you-go Agent API priced from $0.025 per Web Search Agent request at the listed low-effort setting. Companies that want Nimble to configure and manage custom data delivery can instead buy annual managed plans beginning at $2,500 per month.
Where Nimble fits in the emerging agentic search stack
Nimble enters a market that has rapidly expanded beyond traditional web search into autonomous research agents capable of planning, browsing, reasoning, and synthesizing information. Products such as ChatGPT Deep Research, Google Gemini Deep Research, Alibaba’s Tongyi DeepResearch, Perplexity, and Sakana Marlin all seek to automate knowledge work that previously required hours—or, in Marlin’s case, potentially weeks—of human research.
Rather than competing head-to-head as another end-user research assistant, however, Nimble is positioning itself one layer lower in the AI stack—as the web intelligence infrastructure that powers those agents or custom enterprise applications built on leading foundation models.
That distinction reflects an increasingly important architectural shift in enterprise AI. Most “Deep Research” systems optimize the overall research workflow, generating search plans, iteratively gathering information, and producing synthesized reports.
Nimble instead argues that the retrieval layer itself has become the primary bottleneck for enterprise AI deployments. If an agent retrieves too many irrelevant pages or performs unnecessary search iterations, token consumption, latency, and operating costs all increase before the model even begins its main reasoning process.
"Customers across life sciences, insurance, healthcare, pharma, retail, and digital-native companies are all telling us the same thing: we need to feed our agents with more accurate context, and we need to reduce the amount of tokens every task consumes," Knorovich said.
The launch blog makes that argument more concrete by describing how teams frequently rebuild the same retrieval stack themselves. A production agent may start with a search API, then accumulate browser controls, parsers, extraction components, validation steps, memory, caching, evaluations, and custom workflow logic. Nimble is positioning its harness as a managed alternative to that growing engineering burden.
In Nimble’s view, improving retrieval before reasoning begins is more valuable than simply giving a language model more documents to analyze. The company’s Web Search Agents therefore adapt retrieval strategies to a particular workload, combining proprietary indexes with real-time web retrieval and task-specific search policies rather than applying the same search algorithm across every domain.
That makes Nimble less of a direct competitor to OpenAI’s or Google’s research assistants than to developer-focused retrieval infrastructure such as Exa and Tavily. Those platforms also provide AI-native search APIs and research capabilities, but Nimble differentiates itself by emphasizing self-learning retrieval strategies, proprietary indexing, enterprise governance, managed delivery, and token efficiency for production agents.
For organizations building their own AI systems, the distinction could become increasingly important. Foundation models are becoming more capable across the industry, shifting competitive differentiation toward the infrastructure surrounding them—including retrieval, orchestration, memory, observability, and governance. Nimble’s strategy reflects that broader trend, betting that better web intelligence can deliver larger operational gains than incremental improvements in model reasoning alone.
Enterprise infrastructure versus AI research assistants
The different positioning is also reflected in pricing.While consumer-facing AI research assistants are generally sold as productivity subscriptions for individual users or teams, Nimble is pricing its managed service as enterprise infrastructure designed to power production applications. Its pay-as-you-go API, however, gives developers a lower-cost path to test the underlying agent technology before committing to a managed deployment.
Platform
Primary audience
Primary focus
Lowest publicly available price (USD)
Nimble
Developers and enterprises
Managed web retrieval and orchestration infrastructure combining specialized search, browsing, extraction, validation, proprietary indexing, and memory
$0.025 per Agent API request (low-effort setting). Managed service starts at $2,500/month (Startup plan, billed annually).
ChatGPT Deep Research
Professionals, enterprises, and knowledge workers
Autonomous multi-step research with iterative browsing, synthesis, and citations
$20/month (ChatGPT Plus). Higher limits are available with Pro, Team, Enterprise, and Edu plans.
Google Gemini Deep Research
Consumers and enterprises
Research planning integrated with Gemini, Google Search, and Google's productivity ecosystem
$19.99/month (Google AI Pro, U.S.). Higher-capacity AI Ultra and enterprise Workspace offerings are also available.
Tongyi DeepResearch
Developers and AI researchers
Open research model for long-horizon information-seeking and agentic search
Free (open source). Users are responsible for their own infrastructure and cloud compute costs.
Perplexity
Consumers, professionals, and enterprise teams
AI-powered web search and cited research
Free entry tier. Perplexity Pro starts at $20/month with Enterprise Pro available separately.
Exa
Developers and AI platform builders
AI-native search, content retrieval, and asynchronous research agents
Free developer tier (includes monthly credits). Paid Search API pricing starts at approximately $7 per 1,000 requests while Agent runs range from $0.012 to $1.00 per run depending on effort level.
Tavily
Developers building AI agents
Search, extraction, crawling, and research APIs for agents and RAG workflows
Free developer tier (1,000 monthly credits). Pay-as-you-go usage starts at approximately $0.008 per credit.
Sakana Marlin
Enterprises, strategy teams, financial institutions, and research organizations
Ultra Deep Research for hours-long strategic reasoning and executive-grade reports
Pay-as-you-go from approximately $0.61 per credit (¥98/credit) with with 100 credits required per research run (approx $61 per run).
The first subscription tier is Pro at approximately $936/month (¥150,000/month) followed by Team at approximately $2,495/month (¥400,000/month) with Enterprise pricing available by quote.
The comparison reveals three increasingly distinct markets.
ChatGPT Deep Research, Gemini Deep Research, and Perplexity operate primarily as user-facing research assistants.
Exa and Tavily provide developer-facing retrieval and research APIs.
Nimble and Sakana Marlin occupy more enterprise-oriented territory, but at different layers: Nimble supplies retrieval infrastructure, while Marlin performs long-horizon strategic analysis.
Sakana Marlin is particularly useful as a counterpoint. It is positioned as a "Virtual CSO" rather than a search API, running autonomous research loops for as long as eight hours and producing executive-ready reports, references, and supporting materials.
Nimble, by contrast, is designed to sit beneath those kinds of systems, supplying the specialized retrieval, browsing, extraction, validation, and orchestration that enterprise agents need to gather reliable external information before reasoning begins.
The comparison therefore should not be read as a direct price-to-price evaluation. A $20/month ChatGPT Plus or $19.99/month Google AI Pro subscription buys an individual AI workspace with Deep Research capabilities.
Nimble's $2,500/month managed plan funds concurrent production agents, managed ETL, MCP integration, web-page capacity, storage, and hands-free data delivery.
Sakana Marlin's approximately $936/month (¥150,000/month) Pro plan pays for extended, compute-intensive strategic research workflows.
Each price reflects a fundamentally different product boundary and deployment model rather than simply a different level of AI capability.
Why retrieval is becoming the next AI battleground
As enterprise AI systems mature, the industry is increasingly recognizing that model quality alone does not determine application performance.
Large language models frequently fail not because they cannot reason, but because they lack timely, trustworthy external information. That reality has fueled rapid investment across retrieval-augmented generation, AI-native search, web intelligence platforms, knowledge graphs, browser automation, and agent infrastructure.
Nimble’s launch reflects this evolution by focusing less on building another frontier model and more on improving the quality of information flowing into existing ones.
Whether the company’s reported 21-point improvement in answer quality and 51% reduction in token usage hold up across a broad range of enterprise deployments remains to be independently validated.
The larger strategic bet is that, as frontier models become more interchangeable, companies will differentiate themselves through the data, retrieval policies, trusted-source rules, memory systems, and orchestration layers surrounding those models. Nimble is not trying to build the researcher that sits in front of the user. It is trying to become part of the infrastructure that determines what the researcher can find, how efficiently it can find it, and whether the resulting evidence is complete enough to support production decisions.
Web Search Agents are available through Nimble’s API, SDK, and MCP integrations, with a free trial available for developers evaluating the platform.
Target SVP Siobhán Mc Feeney says the AI models her company runs aren't what gives Target its edge — everything built around them is.
"There's a lot in it. That to us is the moat," Mc Feeney said at VB Transform 2026. "The models are great, and they're important. They're just not sufficient to be the competitive advantage."
That discipline shows up early in how Target decides whether to build an agent at all. Mc Feeney was blunt, even "controversial" by her own admission, about the current AI moment: every enterprise wants AI agents, but not everything needs one, she said.
Agents earn their autonomy over time rather than getting it by default, she said — a principle that runs through everything Target has built around them.
Mc Feeney said the goal is to make sure agents are aimed at the problems that drive the most value for Target's guests. “We want to make sure we're investing in the right places," she said.
Being deliberate about agents
Agents are becoming part of Target's underlying architecture, increasingly connecting signals, systems, and decisions across supply chain, replenishment, and demand forecasting.
Mc Feeney framed it as retail's oldest promise — the right product, in the right place, at the right time — delivered at scale.
But her team has been deliberate about building AI agents, beginning with the simplest, most obvious question: What is the problem they’re trying to solve? This leads to several follow-on questions:
Does that problem need an agent?
If it does, what type of agent? An orchestrator? A super agent? A domain-specific agent?
Or is what you're calling an "agent" actually just a tool?
“You define that upfront, and this may sound a little process-heavy, then you have to register and certify your agent,” Mc Feeney said. Because a solution may already exist, and you don’t want to duplicate work.
Agent design kicks off another series of important questions: What triggers an agent to act? Automation? An engineer? A timer? What needs to be put in place to track that?
"We're trying to make sure we have lineage from the very beginning — the birthing of this agent, all the way through — because at 2 a.m. one morning, when something goes sideways, we want to make sure we understand everything that happened," Mc Feeney said.
Autonomy level is another consideration; new agents typically start with base autonomy and earn more over time. What the agent has access to is a separate question: what data, what systems, what tables, what databases?
Finally, there’s monitoring and observability; agents won’t solve problems, or improve over time, if they’re not continuously evaluated.
“We measure everything: What it was intended to do, its calibration, its trajectory, not just runtime and latency,” Mc Feeney said. This creates full transparency, and allows agents to be tweaked over time.
“You're talking about architecture and taxonomy and a data governance layer that absolutely had to be established,” she said.
There's a lot in these "layers of autonomy" — that foundation is what gives Target the ability to scale and properly invest in the right models for the right problem.
Models have different “gradients” that are better for different jobs; for instance, frontier models excel at complex tasks that require crunching billions of pieces of data (like in heavy merchandising supply chains). But in some scenarios they can be cost-prohibitive.
“So it’s making sure there's always a cost benefit,” Mc Feeney said.
Agents must earn their autonomy
A digital-twin simulation predicted men's shorts inventory across three Target stores in Long Beach this summer — and one store came back needing six to seven times more stock than the others, she said. Inventory analysts' first reaction: That can't be right. But the system had found something they hadn't factored in. That store sat less than two miles from the beach; the other two were 10 to 12 miles inland. Analysts let the recommendation stand, and the stock sold through.
"This is science. This is mathematically more significant and more confidence-filling than humans doing it," Mc Feeney said. Results like that are what let Target's agentic systems earn more autonomy over time, she said.
Target looks at AI agent autonomy as "earned" and structures it as a four-level ladder, Mc Feeney said: agents start by making observations without acting, then move to suggesting actions while waiting for approval, then to acting within defined guardrails. At the highest level Target currently operates, agents run end-to-end — but still with a human in the loop.
“The autonomy levels for the agents are super important,” Mc Feeney said. “They earn them, and they can lose them if they don't perform as expected.” Models that drift will be taken out of service.
As she put it, humans earn autonomy when we prove we can do something over time. Nobody is given a bunch of extra responsibilities just because; they have to have shown they’re able to handle them.
In a similar way, agents can be scientifically measured and quantified: how accurate they were, how much they drifted, and how close they came to their intended goal. This helps establish guardrails, allowing builders to work faster, and “go fast forever,” because they're not constantly wondering where the guardrails are.
“If you follow these guardrails, you [follow] security guidelines, you register the agent, and something still goes wrong, we have full lineage all the way through from the start,” Mc Feeney said. “Our ability to recover is much better.”
When it comes down to it, agent success is a confluence of factors, not just one, she said: “It's about your architecture. It's about your taxonomy. It's about the autonomy levels your agents have, and it's about security and observability.”
A new skill set for new workflows
Even when agent autonomy is high, though, builders must still be held accountable when something goes wrong. Mc Feeney noted that teams are now working at speeds no one could have anticipated, which means evaluation harnesses have to be established and agents registered and tracked.
A lot of it is cultural; the workforce is being reshaped and builders and engineers need new skills to manage human workers and AI systems side by side. These contexts are quite different, but the career evolution is “super exciting.”
“You're a builder. You're observing agents building, and you're also coaching humans observing agents building,” Mc Feeney said. “The level of nuance is pretty special.”
Bright Machines wants to solve one of the least glamorous but most consequential problems in the AI buildout: what happens to quality data when a human being has to touch the production line.
The San Francisco-based manufacturer announced today the Hybrid BRC (Bright Robotic Cell), an expansion of its Bright Factory platform that lets human operators step inside a sensor-monitored robotic cell to perform prescribed assembly steps — without breaking the digital record that tracks every server from its first screw to its shipping label.
It sounds like an incremental hardware update. It isn't. The Hybrid BRC is a direct answer to a structural weakness in high-stakes electronics manufacturing — one that CEO Sviat Dulianinov quantified in stark terms in an exclusive interview with VentureBeat.
"If you assemble modern AI servers starting with manual operations, your initial yield — first-pass yield — can be as low as 20%," Dulianinov said. "Then you gradually ramp up and scale, and it can reach the 60s, 65% or so."
When a single AI server can cost hundreds of thousands of dollars, and hyperscalers are burning billions waiting for infrastructure they can't deploy fast enough, that number is the whole story. The Hybrid BRC is Bright Machines' attempt to keep human hands in the loop without letting human error back in the door.
Why manual assembly steps create a black hole in production data
Modern automated assembly lines generate a continuous stream of production data — torque values, placement coordinates, component serial numbers, inspection images. That "data thread" is what lets a manufacturer prove a server was built correctly and, when something fails in the field months later, trace the failure back to a specific station, step, or part.
But automated lines inevitably need manual intervention, and until now manufacturers had two bad options when that happened: stop the line entirely, or pull in-process units off to a separate manual workstation that sits outside the monitored data flow. The first choice kills throughput. The second punches a hole in the production record at precisely the moment when human error is most likely to occur.
The Hybrid BRC eliminates that tradeoff, the company says. The cell incorporates guarded access doors and safety panels directly into the production line. When an operator opens the doors, the robotic arm deactivates, and on-screen instructions guide the operator through each assembly step while the cell's sensor array — cameras, force feedback, and tooling sensors — continues monitoring for incorrect installs, missed steps, and wrong components, applying the same quality checks used during full automation. The traceability record persists at the serial-number level from start to finish.
The yield gap between humans and robots in AI server assembly
The economics driving the design become clear when Dulianinov's manual-assembly figures are set against what automation delivers. "At robotic operations, yield-per-station level is usually more than 98% with our technology, and even at the line level, we usually get to 97.5%, 97.7% or so," he said.
First-pass yield measures the percentage of units that come off the line correct the first time, without rework. The gap between a 20% manual ramp and a 98% automated station isn't a rounding error — it's the difference between profitability and disaster on hardware this expensive.
That math explains the company's design philosophy for the Hybrid BRC, which treats the human operator as an escape valve for exceptions rather than a substitute for automation. "The more human stations you introduce, the more you increase the risk of lower yields driving the overall yield down," Dulianinov said. "That's why we prefer to start at least with 50% automation, and then move to at least 80%." Speed follows a similar pattern: "On the line level, robots can be faster than humans from like 50 to 100%" in throughput terms, he said.
How server assembly became the hidden bottleneck of the AI infrastructure race
The AI infrastructure conversation usually revolves around chip supply, power availability, and data center construction. Dulianinov argues that assembly — the unglamorous work of turning chips and motherboards into racked, tested, deployable compute — is a quietly enormous drag on deployment timelines.
"When you have the chips and you have the motherboards, you want to be as fast as possible to deploy that in the data center," he said, describing greenfield deployments where power and buildings already exist. Getting hardware built, tested, and often rebuilt when quality falls short "could be months," he said. "With more technology used for this, as our tech, we believe that we can cut it by at least a third."
A company executive on the call added an anecdotal but telling data point: the servers Bright Machines produces are "flying out into production" rather than sitting stacked in warehouses awaiting deployment — evidence that assembly capacity, not just chips or power, gates hyperscaler timelines. The stakes are asymmetric, the executive noted, because the largest hyperscalers lose millions of dollars per day when servers fail or arrive late. That is why customers are less interested in buying boxes than in buying assurance — and why an unbroken data thread has become a product in its own right.
Inside the secretive customer base already running hybrid production lines
The Hybrid BRC is not vaporware. Dulianinov said the company already operates a number of the hybrid lines in the U.S. and has "built more than 10,000 compute nodes" through the new stations. This year, he said, Bright Machines plans to manufacture "more than half a gigawatt of compute capacity."
Who's buying? Don't ask. "We cannot unfortunately name customers. That's the toughest part of our job," Dulianinov said. "They're pretty secretive because, as you can imagine, everything data center related is IP related."
He did offer growth figures: customers grew "more than 3x this year" versus the prior year, driven by what he called the intersection of "physical AI, AI infrastructure buildout, and onshoring." The demand is spilling into real estate — the company is moving from its 16th Street San Francisco offices to a Burlingame space this fall that executives described as three to four times larger. Overall, the company says it has deployed more than 130 microfactories across 10-plus countries, served more than 60 customers, and produced more than 300,000 servers.
What separates Bright Machines from Tulip, Instrumental, and contract manufacturing giants
Asked how the Hybrid BRC's traceability claims stack up against operator-guidance and inspection software vendors like Tulip and Instrumental, Dulianinov drew a sharp line around business models.
"Tulip is just a company that does interface for operators. Instrumental, they focus on inspection. It's just pieces of the puzzle," he said. "We, as a technology-enabled manufacturer, we actually run this whole operation... We put our lines, put our software, put our data on the floor, our people, and run it from the beginning to the end."
The right comparison set, he argued, is contract manufacturing giants like Flex, Jabil, and Foxconn — companies that own the full production process but historically built it on manual labor that generates little data. Bright Machines' differentiation, he said, is that robot data, sensor data, and now human-station data all flow through one orchestration layer into a single environment the company calls Bright Insights.
That positioning is notable given the company's origins. Bright Machines was carved out of contract manufacturer Flex eight years ago, and its history has had turbulence: the company planned to go public in 2021 via a SPAC merger at a reported $1.6 billion valuation, according to contemporaneous reporting by The Wall Street Journal and CFO Dive, before the deal fell through. It rebounded in June 2024 with a $126 million Series C — $106 million in equity led by funds managed by BlackRock with participation from Nvidia, Microsoft, Eclipse, Jabil, and Shinhan Securities, plus $20 million in venture debt from J.P. Morgan — bringing its total raised past $400 million, per the company's announcement at the time.
Who owns the production data — and how workers feel about being monitored
For technical decision makers, two governance questions loom over any system that instruments human work this closely, and Dulianinov addressed both directly.
On data ownership, he drew a clean boundary: "Everything related to the customer and inspection of their devices and parts obviously would be protected and owned by the customer." Process and robotics data, he said, stays with Bright Machines to fuel continuous improvement across its platform.
On worker surveillance, he pushed back on the framing. High-IP electronics floors — especially those touching aerospace, defense, or government workloads — already prohibit workers from carrying personal electronics, he noted. "People who know those floors, they know that this is part of the game," he said, adding that employees "actually appreciate" the traceability because it underpins the security mission: "If you build a data center for the government, and then you build servers somewhere in China, you cannot guarantee how exactly it was built and what component was put there." In his telling, the monitoring isn't about watching workers — it's about being able to prove, component by component, that American-built AI infrastructure is what it claims to be.
The onshoring bet: rebuilding American manufacturing without 3 million workers
The Hybrid BRC's modular design carries strategic weight beyond quality assurance. Because the cells are software-defined and snap together like building blocks, Bright Machines says it can retool lines for new hardware generations in days or weeks rather than months — "we can introduce it within a day" for minor design changes within a product family, Dulianinov said, though a jump from air cooling to liquid cooling remains "a big jump." In an industry where new chip architectures now arrive on a roughly annual cadence, changeover speed is arguably as valuable as yield; a production line that takes six months to retool is obsolete before it amortizes.
But Dulianinov's closing argument was about labor arithmetic, not machinery. "We need to build in the U.S., and you don't have 3 million people to bring up manufacturing in the U.S.," he said, referencing the massive workforces of Shenzhen-scale electronics plants. "So you need to solve it with AI software and robots, and that's our thesis... It's not just robots on the floor — it's also creating jobs. All the robots, and some people on the floor."
Lior Susan, founder and CEO of Eclipse and chairman and co-founder of Bright Machines, framed the announcement in the same terms: "The future of manufacturing isn't choosing between automation and flexibility — it's combining both in the same digital production environment."
For all the talk of gigawatts and yield curves, the Hybrid BRC amounts to an admission wrapped in an innovation: even in the most automated factories on Earth, humans still have to open the door and reach inside. Bright Machines' wager is that the winners of the AI infrastructure race won't be the manufacturers who eliminate the human hand — but the ones who never lose sight of it.
Visa aimed Anthropic's Claude Mythos at the infrastructure behind billions of daily transactions, a network that spans more than 200 countries and territories, moves money in roughly 160 currencies, and connects nearly 5 billion payment credentials to more than 175 million merchant locations.
The model stitched minor weaknesses deep in the stack into working exploit chains that would traditionally have surfaced only late in penetration testing. Rajat Taneja, Visa's president of technology, walked the VB Transform 2026 audience through what came next, including why Visa released the harness that governed the entire hunt as open source and why the company abandoned traditional remediation metrics for a measurement its team invented.
Taneja has run technology strategy, product engineering, and global infrastructure at Visa since 2019, after joining the company in 2013 from Electronic Arts, where he served as CTO following 15 years at Microsoft. He co-authored, with Visa chief information security officer Subra Kumaraswamy, the June 10 blog post announcing the release of the Visa Vulnerability Agentic Harness on GitHub as a reference implementation that any security team can inspect, adapt, and extend. Visa also published a technical white paper detailing the architecture, lessons learned, and 12 non-negotiable architectural practices for critical infrastructure.
Trust built on pessimism and paranoia
Taneja led with the arithmetic that makes Visa a target worth defending obsessively. Trust at the scale of global payments gets engineered through what he called pessimism and paranoia, by assuming failure and designing around it before failure arrives. The network has been hardened over many years through zero-trust architecture, layered defenses, and highly automated security operations built for the scale and reliability global payments demand.
So when Anthropic invited the organizations behind critical software to test Mythos under Project Glasswing, Visa said yes. Glasswing participants collectively identified more than 10,000 high- or critical-severity vulnerabilities in the first month of testing across software underpinning critical systems industry-wide, according to Anthropic. Anthropic's own conclusion placed the bottleneck after discovery, in verification, disclosure, and patching speed. Visa joined to test decades of hardening at AI speed and learn where advanced models could push its defenses further.
What Mythos showed at Visa
Inside Visa's environment, Mythos demonstrated system-wide, context-aware analysis, surfacing vulnerabilities buried deep in the stack and flagging issues that grow more serious when chained together, with findings clean enough that engineering teams could act on them without wading through noise. Some findings carried critical severity ratings, and Visa credits its zero-trust controls, network segmentation, and layered safeguards with breaking the chain before any external actor could have acted.
That confirmation mattered, Taneja said, but the epiphany that followed mattered more. "In a world of agentic attacks, defense also has to be agentic," he said. Even at a company that has invested decades in defense-in-depth, the model revealed assumptions the team had been operating under that needed rethinking. Traditional SAST tools keep their place as a first pass against known vulnerability patterns, Visa's white paper notes, but pattern matching alone cannot follow an adversary who reasons through logic, data flow, and the exploit chains that live between the signatures.
A harness, not a scanner
Visa's response was not another monolithic scanner. The team built the Visa Vulnerability Agentic Harness, now in its fifth generation, as a governed pipeline that directs frontier AI models through structured security tasks while enforcing deterministic controls, policy gates, and human oversight at every stage. Taneja walked through the design philosophy. The harness operates across four phases and eleven stages, from code ingestion and threat modeling through deep-dive verification, exploit chain synthesis, and finally remediation and fix validation.
Three design choices drive finding quality, per the project's own documentation. Threat modeling runs before analysis to focus on the attack surface rather than scanning everything blindly, multi-agent deterministic voting requires convergence across independent reasoning chains before a finding advances, and structured triage artifacts compress the lifecycle from discovery to a result developers can actually ship. The payoff is a pipeline that runs hot by default. A plain scan in the shipped profile runs all eleven stages and edits source files in the target repository in fix mode, applying candidate patches unless the operator stops it at detection.
The harness is multi-model by design. An LLM abstraction layer lets Visa swap or combine providers without changing the control plane, and the open-source version works with Anthropic Claude, OpenAI-compatible models, or a mix. The repo's documentation is candid about the exception. Applying a fix requires the file-editing tools that only the Anthropic backends expose, so the remediation and validation stages currently require Anthropic models for full functionality, and an OpenAI-compatible model in those roles is limited to report-only output. VentureBeat's Q2 2026 Pulse research, presented earlier at the conference, reinforces why that provider flexibility matters. Among the enterprises surveyed, 82% rely on provider-native controls as their primary security layer, and 59% plan to adopt or switch agent security tooling within the year. The controls enterprises adopted last year are already becoming the controls they plan to replace.
Mean Time to Adapt replaces legacy metrics
Finding vulnerabilities is no longer the hard part, Taneja argued. The real challenge is how quickly a team can confirm an issue is truly exploitable, fix it, and prove the attack path is closed rather than just showing a patch was applied. Visa calls this Mean Time to Adapt, and the white paper tracks it along three dimensions. Inventory freshness measures how current and complete the organization's view is of code, configuration, and runtime deployment. Exploitable paths per release counts how many end-to-end attack chains remain possible after each release, not just how many findings were closed. Validation cycle time tracks how long it takes to produce repeatable, evidence-backed proof that a fix works and stays working in production.
That distinction matters because legacy measures such as mean time to detect and raw CVE closure counts can look better on paper while actual exposure keeps growing underneath them. An organization can close hundreds of findings a month and still leave viable exploit chains open if nobody tested whether the patches actually break the attack. MTTA forces teams to measure the outcome that matters, and the white paper leans on CISA Known Exploited Vulnerabilities data to make the prioritization case, noting that fewer than 1% of CVEs are ever actively exploited. Visa's SSDLC policy now assumes every exploitable path will be exercised in production and requires it to be remediated before code is promoted.
Supply chain risk accelerates under AI
The conversation moved past Visa's own perimeter when Taneja turned to suppliers. A well-defended enterprise stays exposed through weak vendors and weak open-source components, the white paper warns, so Visa is making AI-specific security posture a non-negotiable dimension of supplier due diligence, with expectations for continuous vulnerability validation, living software bills of materials, and MTTA baselines across its technology stack.
Visa has also joined Project Lightwell, the $5 billion IBM and Red Hat initiative to harden widely used open-source components through AI-driven validation and coordinated patching, alongside financial institutions including Bank of America, JPMorganChase, Goldman Sachs, and Mastercard. The commitment extends the same logic upstream, because the MTTA clock does not pause at any single company's perimeter.
When agents start buying things
Securing agentic commerce is Visa's next problem. Taneja described a future where AI agents transact on behalf of consumers and enterprises, and said Visa is building the trust framework, identity layer, and agent readiness scoring that merchants will need before agents can safely complete transactions. Behind that work sits the Visa Payment Threats Lab, a simulation environment where real fraud scenarios get replayed against the authorization rules, thresholds, and configurations Visa actually runs, to surface AI-enabled failure modes as targeted hardening recommendations.
The identity challenge is not theoretical. VentureBeat's Pulse research found that 69% of enterprises already run credential sharing somewhere in their agent deployments, and companies with shared credentials report security incidents or near-misses at a 63.5% rate, against 40.9% where every agent has its own scoped identity. Visa's white paper addresses that gap directly, listing "AI agents are identities" among its 12 non-negotiable practices and requiring scoped permissions, least privilege enforcement, full audit trails, and inclusion in IAM governance for every agent that calls an API, reads data, or modifies a system.
Three priorities for defenders
Visa is organizing its defensive strategy around three priorities, Taneja said. Shift security left until exploitable flaws are designed out before they reach production, and replace high-risk, under-supported components before they turn into material exposure. The third is the heaviest lift at Visa's scale, refactoring defenses to run autonomously under human governance so detection, validation, and response keep pace as threat volume grows and the models behind attacks improve.
None of it requires a payment network's budget to start. The harness sits on GitHub with 595 stars and 97 forks as of July 20, MTTA needs a dashboard rather than a procurement cycle, and the white paper's 12 non-negotiable practices map onto architecture reviews security teams already run. Visa's own conclusion reads like a deadline. The opening to get ahead of machine-speed attackers is still there, the paper argues, and it will not stay open.
Instacart is posing the provocative question: What if most of the work your engineers do today should, in fact, be done by machines?
At VB Transform 2026, CTO Anirban Kundu argued that dev teams continue to waste their time on draining, repetitive, high-volume work; this should be absorbed by AI agents so that humans can focus on problems that require judgment, intent, and exception handling.
In fact, in 97% of cases, Instacart’s builders don’t even read code anymore.
“In the past, the tactical level was the creation of the code,” Kundu said. “In the most tactical level going forward, it's going to be, ‘How do you navigate around the AI system to give you what you want?’”
AI generating code, performing "pretty serious evals"
That doesn’t mean humans never look at code; agents handle the bulk of code generation and boilerplate, particularly with newer projects where code is generated or regenerated on a weekly basis.
“The benefit of that is we don't care about tech debt anymore,” Kundu said. “Things that are not active just get dropped out and then it gets rebuilt, kind of like how we used to build assembly code or object code.”
So why not 100%? The remaining 3% is in legacy, compliance, and latency-sensitive systems and workflows, or driven by a “boatload of code” that is dead, not active, or half-active. These cases still need careful human attention.
Instacart is slowly “smoothing those parts out,” however, breaking systems down in an aptly-named project Atoms, then building them back up in a cleaner, more modular form. Kundu’s team started with the “monoliths” and is shifting to remote procedure call (RPC)-driven architectures.
But evaluation remains one of the overarching challenges. Code reviews aren’t as relevant when AI is generating code — as Kundu noted, “the lines of code are going to be correct, the syntax is going to meet your expectations” — so the goal is to move to an “intent model.” That is, training devs so they can ask different models the right questions from an intent perspective.
Evals are then performed independently: Roughly 7,000 automatic evaluations run each month, and the system answers 8,000-plus real-time developer queries with about 99.9% accuracy.
Identifying "hiccups" that human intuition might have missed
Dovetailing with this, Instacart has built an agentic site reliability engineering (SRE) system trained on years of the company’s own incidents and root-cause analyses rather than generic failure data. Instead of teaching a model how production outages work in the abstract, the team fed it the specific ways Instacart’s systems have broken over time, along with the ways humans diagnosed and fixed them.
As a result, the company has seen accuracy in detecting and mitigating production issues jump from roughly 60 to more than 90%.
Kundu pointed to one example with Instacart’s internal tool Blueberry. The AI SRE colleague watches 200-some-odd Slack channels, monitors signals, and looks for patterns across human conversations and alerts.
In one incident, a database shard backed by an EBS volume that had a “hiccup” for a period of time. The human team did not immediately suspect AWS disk issues and were “obviously scrambling” to figure out why this particular shard misbehaved.
But about 20 minutes in, Blueberry posted on Slack, pointing to a specific blip and tying it to a feature-flag-like system called "roulette" that had been inadequate. "It's supposed to be rolling out in this cadence, [but] it had been too much,” Kundu said.
Blueberry figured it out, and the team resolved the incident. “Would have a human been as quick? I think the problem is human intuition would hold us back a little bit,” Kundu said.
Humans tend to default to patterns we’ve seen before, then resort to debugging; Kundu called this the “first brain-second brain kind of thing.” But Instacart’s agentic SRE is actually “more comprehensive in its ability to look at everything and then be able to decide what does or doesn't matter.”
Redefining the engineer’s job
Looking ahead, the most tactical work for engineers will be navigating AI systems: Designing and supervising evaluation processes; coordinating multiple simultaneous experiments and features; managing constraints like limited top-of-funnel traffic for testing; figuring out when to escalate; identifying edge cases and where things might break.
Domain expertise is also being rethought in the age of AI. Instead of bottlenecking changes through a single “owner” team that touches the code, Instacart is embedding domain knowledge into definitions and specs that any team can use.
“We’ve lived in this world where this group or this engineering team is the one that can touch the code and make the modification,” said Kundu. “We're trying to move into a world where the code becomes completely democratized across groups.”