❌

Normal view

Teradyne invests in Bright Machines to advance AI infrastructure manufacturing

5 October 2026 at 19:46
Teradyne, a provider of automated test equipment and advanced robotics systems, and Bright Machines, a next-generation manufacturer bringing AI and data center infrastructure production to the edge, have announced a strategic investment by Teradyne in Bright Machines. The investment accompanies a strategic collaboration focused on integrating Teradyne robotics and test technologies with Bright Machines’ software-defined […]

Anthropic’s answer to Dots and Muse is already inside Claude

I’m Matt Burns, Chief Content Officer at Insight Media Group. Each week, I round up the most important AI developments, explaining what they mean for people and organizations putting this technology to work. The thesis is simple: workers who learn to use AI will define the next era of their industries, and this newsletter is here to help you be one of them.


OpenAI launched Dots at DevDay on Tuesday. Each Dot is an always-on agent with its own cloud computer and browser, hooked into the 4,000-plus apps that already connect to ChatGPT. In OpenAI’s own example, a Dot sees a bug alert land in Slack and starts digging in on its own. Cool.

And before Dots, Meta launched Muse, and it’s crushing the mobile install numbers previously set by ChatGPT. And before Muse, xAI shipped its version, Grok Bot, in August. 

Anthropic’s version, though different in a couple of ways, arrived two weeks ago as an update to Claude, and without a cute name or fuzzy mascot. This update folds Cowork, which has worked since July to run scheduled jobs after a user closes their laptop without being asked, into the main Claude app.

Anthropic made the right call with Cowork, whether or not it ever matches Muse’s downloads. Always-on agents are too young to have a winner, and the moats are shallow. A feature one lab ships tends to show up at its rivals within weeks or months. Meta needs Muse to be a blockbuster. Anthropic needs the people already building with Claude to hand it real, recurring work and keep coming back. Repeat use and finished jobs matter more than a flashy launch.

Always-on agents are too early to have a winner

This wave of always-on personal agents took off less than a year ago. Peter Steinberger pushed a weekend project called Clawdbot to GitHub last November. It lived on a spare computer, took prompts over WhatsApp or Telegram, and relied mostly on Claude. Anthropic sent the lawyers in January, so it became Moltbot, then OpenClaw a couple of days later. It turned into a security headache, and in February Steinberger joined OpenAI. On Tuesday, the foundation that runs the project released an early, pre-1.0 version of OpenClaw Enterprise for companies to try internally.

Ten months, three names, one OpenAI hire and an enterprise edition. Ideas move between these products faster than ever. Meta’s Nat Friedman said Muse was heavily inspired by OpenClaw, and users digging through Muse found a SOUL.md personality file nearly identical to OpenClaw’s.

A feature one lab ships shows up at its rivals within months.

Where an early version of each AI assistant and agent feature launched, and who offers one now.

Feature Example of an early launch Now also at
Deep research reports Google Gemini (Dec. 2024) OpenAI, Anthropic, Perplexity
Terminal coding agent Claude Code (Feb. 2025) OpenAI Codex, Gemini CLI, Meta Muse Code
Agent personality file OpenClaw (SOUL.md) Meta Muse
Agent with its own work identity Anthropic Claude Tag (June 2026) OpenAI specialist Dots (preview)
Chat and agent work in one window OpenAI Work mode (July 2026) Claude (Sept. 2026)

Sources: company announcements; The Next Web (SOUL.md); OpenAI (specialist Dots); The New Stack (Work mode and the Claude merge). Dates mark an early example, not necessarily the first.

Anthropic’s September merge is on the list above as a copier, two months behind OpenAI’s Work mode — though Cowork’s cloud and scheduled-task features shipped July 7, two days before Work mode. Jessica Wachtel tested both on our site across three developer tasks and found them tied on accuracy, with ChatGPT faster and Claude more thorough. When a product is this young, a me-too launch is how a lab gets its own users in the room to see what they do with it.

Meta needs Muse to be huge. The others need paying users.

Muse is a hit. Counting only iOS, Apptopia estimated 359,000 daily U.S. users in Muse’s first 12 days, compared with 231,000 for ChatGPT’s iOS app at the same point. Sensor Tower estimated more than 3.4 million downloads as of September 24, though other firms’ counts vary.

Meta needs this. Its Reality Labs division spent tens of billions of dollars on the metaverse without a mainstream product to show for it. Llama 4 landed so flat last year that Zuckerberg rebuilt his AI division around a new superintelligence lab, notably taking a 49% stake in Scale AI and hiring its founder Alexandr Wang. 

Muse is the first product from that rebuild to break through, and it got there on Meta’s user base: Apptopia found that more than 95% of Muse users also use Facebook. Muse has a free tier, paid plans at $20 and $100 a month, and connectors to Shopify and Stripe checkout, enabling agents to buy things. Meta’s route runs through scale. Scale has a downside, too. Janakiram MSV explains for TNS why Amazon started blocking Muse two weeks after launch, and it’s well worth your time because this is a new wedge in ecommerce.

OpenAI and xAI started at the paid end. The price of entry for Dots? A ChatGPT Pro plan at $100 to $500 a month, or a Business Premium or Enterprise account. Sam Altman called Dots a premium product because each one needs so much compute. Grok Bot reached 418,000 weekly users by mid-September, per Bloomberg, and xAI just launched Team Bots on Monday.

Neither company needs Meta’s install base to learn something useful. Paying customers already spend money on AI, and retention will show whether they keep using the agent once the novelty wears off. 

Anthropic shipped its version to the people already building with Claude

Anthropic did what I’d expect from labs right now. It shipped fast, shipped to paying subs, and put the agent inside the app those people already use. The merged Claude hit Pro and Max customers, with Team and Free plans coming soon. Teams have had Claude Tag in Slack since June, a proactive teammate with its own identity and audit trail. Cat Wu said an internal version accounts for about 65% of the product teams’ code changes. Developers have had Managed Agents for hosted, long-running agents since April.

What Anthropic hasn’t done is put Claude where Muse lives. Muse works inside Meta-owned WhatsApp, but Claude still asks you to open the Claude app or tag it in Slack. That workflow might matter for mainstream consumers. Teams wiring an agent into their code care more about what it can touch. Jani laid out the difference on our site in August: Grok Bots share one cloud computer and one set of logins across the whole roster, while Claude Tag joins a Slack workspace with its own service identity and channel-scoped access. Before handing agents your logins, know what you’re getting into.

Those are the users Anthropic should want. Writers on Towards Data Science have spent the year showing what builders do with always-on agents. Samir Saci put a team of OpenClaw agents on a supply chain simulation to chase down late shipments. Eivind Kjosbakken’s guide to running a fleet of OpenClaw bots walks through nightly QA bots that test an app and report bugs, plus agents that check invoices. Both ran their agents on OpenAI’s Codex. OpenAI runs its own OpenClaw agent, Androidclaw, that traces broken builds and, in some cases, merges the fix, VentureBeat reported. Anthropic needs that kind of work running on Claude.

I set up OpenClaw on an old Mac Mini this spring, and I had so many questions (and breakthroughs). That’s why builders, coders, and developers are critical to product development: They ask these questions out loud in GitHub issues and Discord threads, and the answers end up in the product.

The obvious objection is trust. It’s fair. An always-on agent holds credentials and acts while nobody is watching. xAI’s own documentation tells users not to treat Grok Bots as a security boundary, and the OpenClaw Foundation says most IT departments ban agent platforms outright. Anthropic’s defaults lean cautious: Claude asks before it acts unless you tell it otherwise, and Tag keeps its own audit trail. If builders don’t come back, or each finished task takes more supervision than doing it themselves, the experiment hasn’t worked. But if they keep finding useful work to hand off, their questions and breakthroughs become the roadmap.

That’s the race that matters for Claude, Dots and Muse: becoming the agent people trust with the next job.

The post Anthropic’s answer to Dots and Muse is already inside Claude appeared first on The New Stack.

AMD to acquire World Labs in deal worth $8.2 billion

1 October 2026 at 17:00
Chipmaker AMD has entered into a definitive agreement to acquire World Labs, an AI model and research lab led by AI pioneer Dr. Fei-Fei Li. The acquisition will bring a world-class team of researchers and model experts to AMD, strengthening its ability to develop AI hardware, software and systems around the needs of emerging models […]

Solomon Asamoah on Ghana’s AI Infrastructure Test: Compute Cannot Outrun Power and Water

30 September 2026 at 18:16
Artificial intelligence is usually discussed as software, but its limits are increasingly physical. Every model depends on processors housed in data centres, continuous electricity, cooling systems, fibre connections and engineers who can keep the equipment running. Solomon Asamoah sees this as a defining infrastructure question for Ghana and Africa: how can the continent expand compute […]

Microsoft Fabric is where AI agents learn how the business works

Microsoft is turning Fabric, its integrated data platform, into the place enterprise agents go to learn how a company works, whether that’s what happened last quarter or what’s happening on the factory floor right now.

On Tuesday, at its FabCon and SQLCon data conference in Barcelona, Spain, Microsoft outlined the next step in its vision for how the products that make up Fabric can simplify all of this information gathering.

Microsoft, for example, described how Fabric IQ, the context layer inside Fabric, now feeds Microsoft 365 Copilot by default. Outside agents can query it over MCP. Power BI will soon be able to turn the same definitions into apps. And ontologies, which add business rules to the mix, moved further into preview.

Credit: The New Stack

As Microsoft Fabric CTO Amir Netz said in a press briefing after the keynote, “Agents are a very, very strange animal, and I always like to say it’s like Drew Barrymore from 50 First Dates. Every time they open their eyes, they forget everything that happened before.

“They don’t know where they are, so the first thing we have to do is tell them where they are. You are now working for Microsoft. You are now working for Wells Fargo. You are now working for Emirates.”

“Agents are a very, very strange animal, and I always like to say it’s like Drew Barrymore from 50 First Dates. Every time they open their eyes, they forget everything that happened before.”

Customers tend to find this out the hard way. Yitzhak Kesselman, the corporate vice president who runs Fabric IQ, tells The New Stack that companies unify their data, run models on it, and then look at the answers.

“Customers that are more advanced on their journey have their own evals for their agent,” he says, adding that what they see sometimes isn’t what they expected. “Then [customers] will understand: ‘Okay, now I need to create the context for my agents.'”

Kesselman says he met with more than 320 companies last year, and the pressure to get there comes from the business side, which asks to “‘show me the value of those agents,'” he says, “‘the before and after.'”

Arun Ulag, Microsoft’s executive vice president for Azure Data, said in the keynote that coding agents work because they have “the code, the repos, the change history, the specs, the tests.” But outside of coding, most enterprises have nothing comparable, he argued.

Fabric IQ is one of four pieces of what Microsoft calls Microsoft IQ (this is Microsoft, after all, so there’s always a lot of different names and products involved).

Work IQ covers email, Teams, and SharePoint. Foundry IQ covers documents and manuals. Web IQ covers the public internet.

Fabric IQ, Ulag said, “focuses on the state of your business and how your business actually runs.” His announcement blog post describes it as combining “unified data from OneLake, trusted metrics from Power BI semantic models, and operational context from ontologies and real-time intelligence.”

Getting the data in

“Everyone wants to jump to AI, but then they realize, ‘Oh, I need data for it,'” Kesselman says. “They need data from their business applications, their structured data. They need to have real-time data.”

OneLake, Fabric’s storage layer, is at the core of all of this, and for the system to provide context, enterprises have to feed it data from across all the first-party and third-party services they use.

Fabric’s shortcuts and mirroring, which connect or replicate external sources into OneLake, are free of charge, and Netz said storage under management in OneLake is growing 300% year over year.

What’s new this week at FabCon and SQLCon Barcelona 2026

On Tuesday, the company announced a two-way integration with Salesforce Data Cloud 360, mirroring from Google BigQuery in general availability, a ClickHouse workload that runs that engine directly against OneLake, and dbt’s Fusion engine, due in Fabric in the next couple of weeks.

Simply copying data won’t satisfy security teams, though. Mirrored security roles, which reach public preview in the coming weeks, bring permissions along with the data. Access roles defined in Snowflake, for example, will show up in Fabric with the same members and the same table permissions.

As Microsoft’s Shireen Bahadur, who demonstrated the feature in the keynote, put it, “We’re not just solely replicating security rules. We’re actually actively enforcing and preserving those roles.”

“We’re not just solely replicating security rules. We’re actually actively enforcing and preserving those roles.”

OneLake isn’t a one-way street. Enterprises can also pull data out and use it in other tools.

OneLake data, after all, is stored in open formats and has open APIs, and Netz said on stage that any engine that reads them can use it. “If you want to use this data that we help people to get for free with our competitors’ tools, go for it. I’m not happy about it, but go for it,” he said.

IQ sharing, now in preview, lets an organization share governed tables, files, Markdown agent instructions, and RDF ontologies with another tenant without copying, with an expiration date. That’s something Fabric users have wanted for a while, and the keynote announcement drew plenty of applause from the thousands of data professionals in attendance.

Past, present, and future

“Not only what we have in the past, not only what’s happening in the present. We also want the agents to understand what we want to happen in the future,” Netz said in the keynote.

The past is semantic models. A semantic model is the layer under every Power BI report that defines how a metric like revenue is calculated, how entities relate, and which tables supply the numbers. Netz said Power BI users have already created 22 million, and Ulag called them “the core of Fabric IQ.”

The present is real-time intelligence, Fabric’s streaming and event stack, which Kesselman also runs. He recalls two CIOs telling him the same thing that day in Paris. “‘There is no AI without RTI;’ there’s no AI without real-time intelligence,” he says. “If you don’t have the streaming data… data that is fresh, you will run your LLMs on data which is like hours or days old.”

“‘There is no AI without RTI;’ there’s no AI without real-time intelligence.”

A signal on its own isn’t enough, though: “You want to take the signal that you get now… the event that comes now, and look historically, okay, is it an anomaly or not?” Kesselman says. That’s why the keynote showed batch copy jobs and event streams feeding each other, so an agent can judge what just happened against what usually happens.

The future is Fabric Planning, which went generally available in July and got a performance and feature update this week. It’s Microsoft’s tool for budgets, forecasts, and targets, the numbers a business wants to hit rather than the ones it has already recorded.

A plan is a Fabric item like a lakehouse or a notebook, and it borrows its measures from the same semantic models the reports use, so revenue means the same thing in the forecast as it does in the dashboard. In the demo, a change to the inputs cascaded through a model of about 13 million cells.

What makes these work is ontologies: Formal descriptions of the things a business deals with, how they relate, and the rules for handling them. Netz called them “semantic models plus plus” on stage, and in a later press briefing, he used an airline to explain the plus.

“If you are an airline, you have planes, you have pilots, you have ground crews, you have airports, you have luggage,” he said. Those are the entities.

The relationships between them “are way more than data relationships,” he said. “It’s not only to say, ‘Oh, I have a match with a foreign key and a primary key between a pilot and a plane,’ but it could be a policy relationship. Which pilot can fly which plane based on their certification, or is the pilot allowed to fly this plane right now based on the rest hours in the last 24 hours?”

But ontologies don’t only define nouns; they also define verbs. “I can ground a plane. I can redirect the plane. I can assign a plane a gate.”

Creating those ontologies, which have to map the various terms that a company might use for the same thing to a single entity, can be extremely time-consuming. But since they are so core to this project, and to allowing agents to reason over data, Microsoft built a tool to generate them automatically.

Putting it to work

Fabric IQ in Microsoft 365 Copilot Chat and Cowork, the mode for delegating multistep tasks that Microsoft built on the technology behind Anthropic’s Claude Cowork, is generally available as of Tuesday. It answers business questions from Power BI semantic models and reports; it’s on by default for Fabric and Power BI customers, and Microsoft says it doesn’t consume additional AI tokens.

Credit: The New Stack

But business users don’t just want to chat with these models. In this age of vibecoding, they also want to generate applications that use all of this data.

Power BI Desktop is getting an app-creation experience in preview in the coming weeks. Users start from a semantic model, describe the app, and then have Copilot generate and publish it to Fabric.

Unlike with a report, Ulag wrote, these apps “can accept inputs, write back data, preserve shared state, and support operational workflows.”

They’re Fabric Apps built on Rayfin, the open-source SDK Microsoft introduced at Build. Power BI Pro and Premium Per User customers get them with a Fabric database of up to 1GB per app at no additional cost.

Netz confirmed in the briefing that “there’s no core difference” between the two, and Ulag added, “What gets created is a Fabric app, right? And that’s it.”

For agents outside Copilot, Fabric IQ MCP is generally available now, with six read-only tools for finding semantic models and reports, reading their schemas, and running DAX queries against them.

Ontology MCP tools, currently in preview, expose ontology definitions and queries. Fabric data agents can now, also in preview, use an ontology as their context source. And the context reaches Microsoft Foundry, Copilot Studio, and GitHub Copilot through the same layer.

Developers can drive the new data engineering agent from their own tools.

Built on the Osmos technology Microsoft acquired in January, it takes on long-running work like migrations and ETL, and can be started and steered from GitHub Copilot, VS Code, Codex, and Claude Code.

Frontier models will write that code, Netz said on stage, and “it won’t fail,” but you don’t know whether the results are right.

Everyone has a context layer now

This year, Databricks shipped Genie Ontology, Snowflake shipped a Horizon Context Layer, Google shipped a Knowledge Catalog, and Salesforce shipped a headless Data 360, which it describes as a context service. Databricks CEO Ali Ghodsi summed up the pitch at his own conference in June as AI having a context problem rather than an intelligence problem.

Microsoft’s version turns on by default inside Microsoft 365 Copilot and starts from 22 million semantic models. Microsoft also said this week it will contribute to Apache Ossie, the vendor-neutral semantic metadata standard Snowflake started, and that it wants DAX, Power BI’s formula language, recognized as an Ossie query language.

When Microsoft made the context argument at Build in June, the pieces that let an agent act on that context were mostly on a roadmap. Now the answers are generally available, and the actions are in preview.

Asked how Fabric copes when the user is no longer one person running a dashboard but a fleet of agents, Kesselman points to observability.

“Fabric really allows you to have this kind of holy grail of combination of both kinds of system data, how the agent runs, but also the business data,” he says, so a company can check whether an agent did what it was supposed to do and what that did to the business. “There’s more in this area that will come in the next few months.”

The post Microsoft Fabric is where AI agents learn how the business works appeared first on The New Stack.

Restate lands $20M as the need for durable infrastructure increases with AI agents

Instead of building its durable execution engine on top of an external database, the company developed its own storage, replication, and redundancy layers. This architecture allows Restate to be exceptionally fast and lightweight.

Eclipse wants companies to be free to switch AI providers. Today, doing so can mean a costly rebuild.

Abstract 3D render of a low-poly wireframe mesh in purple, teal and green, with translucent triangular facets connected across a dark blue background.

If you rely too heavily on one AI provider, switching can mean rebuilding workloads and moving data while you face untangling integrations. A new vendor-neutral group aims to help organizations keep their options open. That’s the goal of the Eclipse Foundation, which announced the formation of the Sovereign AI Foundation on Wednesday.

The new body will serve as a vendor-neutral initiative to help member organizations manage AI platform lock-in, monitor cyber risk, and maintain control of their data. The group has been active on Eclipse’s project site since July.

An initial community of 17 organizations comes together to advance AI sovereignty through open-source collaboration.

The foundation’s mission is to help organizations “navigate AI dependencies” (as in dependence on any single AI service, not control over where application dependencies get tied to a particular model’s DNA… although, by extension, that too), where dependence can limit a software engineering team’s ability to control where data is processed, choose which technologies to use, or change direction as costs, requirements, and provider terms evolve.

A core rationale for AI sovereignty control

Executive director at the Eclipse Foundation, Mike Milinkovich, tells The New Stack that the risk of not adopting an open approach to AI sovereignty control depends on how deeply AI is embedded in an organization and how difficult the underlying systems are to change.

“The deeper the dependency, the greater the potential risk,” Milinkovich says. “But however embedded the underlying systems are, the costs show up in predictable places: higher prices that make switching difficult, and the challenge of rebuilding integrations, moving data and workloads, documenting systems after the fact, compliance delays, and operational disruption.”

He underlines his point and says part of the Sovereign AI Foundation’s work will be to “compare real-world experiences” and build stronger evidence around those risks.

“The deeper the dependency, the greater the potential risk. But however embedded the underlying systems are, the costs show up in predictable places.”

The Sovereign AI Foundation will work to give members an open forum to understand AI dependencies, compare experiences, and develop practical guidance for decisions about data, models, infrastructure, and operations.

Can we move beyond the AI hype-cycle now?

Acknowledging what he calls the pre-IPO hype cycle around the promotion of frontier models, Milinkovich calls for a concerted focus on where the real value in AI lies: enterprise and industrial applications that deliver measurable business value, without vendor lock-in.

“AI will not deliver on its potential if it is based entirely on a dependence on a relatively small number of providers for critical AI models, platforms, and infrastructure,” Milinkovich insists. “As AI becomes part of the systems organizations rely on every day, that dependence can leave them with too little visibility into where their data goes, which models use it, or how easily they can switch when costs, terms, or requirements change.”

Choice, control &  freedom of action

He asks us to understand how far the Sovereign AI Foundation will go to find practical ways to “preserve choice and control,” with a vision of delivering open-source and open-weight solutions that allow freedom of action.

As admirable and philanthropic as this sounds — and coming from a Brussels-based foundation whose founding members are mostly European — is there any America-first idealism here, given that open-weight Chinese models are also openly available?

“This is not about China or any single provider. The same questions apply to AI technologies from the United States, Europe, Asia, or anywhere else,” confirms Milinkovich. “Closed models can limit visibility, while access to model weights alone does not provide the transparency, control, or the independent governance that organizations may need. The real question for our members is whether they can understand, adapt, and, if necessary, replace the systems they depend on.”

17 founding members

As noted above, this initiative launches with 17 participating organizations spanning technology, industry, research, and the open source community: France’s CEA LIST,  Thales and Kentyou, Germany’s EclipseSource, Open Elements, TypeFox, Vector Informatik, and Bosch, Italy’s Engineering Ingegneria Informatica and Eurotech, Sweden’s Ericsson, India’s Infosys, Belgium’s KU Leuven, Japan’s Renesas Electronics, Malaysia’s The IO Foundation… and then there’s North Carolina’s Red Hat, and the University of York in the UK.

The Sovereign AI Foundation focuses on five areas:

  • Mapping real-world use cases: Building a shared view of how members deploy, evaluate, and pilot open-source AI.
  • Assessing technologies and trends: Monitoring developments across models, evaluation frameworks, inference infrastructure, and regulation.
  • Developing practical resources: Producing whitepapers, blueprints, landscape analyses, and reference materials.
  • Building a peer community: Connecting organizations at every stage of open source AI adoption.
  • Translating research into action: Turning outcomes from Eclipse Foundation research initiatives, including EU-funded projects, into insights that members can apply.

A mission to tackle data control, cyber safety, or lock-in?

There’s a lot to unpack here in an organization only just out of the blocks; it begs the question whether the Sovereign AI Foundation will focus on data control challenges and efficiency, cyber safety, or platform lock-in challenges.

“It’s all three, and they reinforce one another,” Milinkovich says. “It is an efficiency issue if moving to a better or less expensive model requires rebuilding the workload. It is a security issue if organizations lack the visibility needed to assess risk. And it is a lock-in and resilience issue if critical operations become so tied to one provider that changing direction is costly or disruptive.”

“The need here is clearest wherever AI is used in systems with long lifecycles, demanding safety and security requirements, or sensitive data.”

He explains that “the specific needs vary”, so a manufacturer may need AI to keep operating when connectivity is limited, while a public agency may need to explain how an AI-supported decision was made.

“The need here is clearest wherever AI is used in systems with long lifecycles, demanding safety and security requirements, or sensitive data. That includes automotive, manufacturing, telecommunications, aerospace and defense, energy, and the public sector,” adds Milinkovich.

Interest will extend well beyond Europe

He confirms that the foundation’s initial participants do “reflect that mix”, as companies including Bosch, Vector, Renesas, Ericsson, and Thales put sovereignty high on the agenda. But he qualifies that dependence on critical AI technologies is a global concern, “so we expect interest well beyond Europe” in the weeks and months ahead.

Milinkovich offers a selection of still-forming use cases that he thinks are already solid. 

  • EclipseSource is working with industrial customers on AI-enabled development tools like Eclipse Theia AI, where teams need flexibility over models and deployment.
  • The Open VSX Registry provides vendor-neutral access to thousands of extensions for a growing range of open source and AI-enabled development tools.
  • Eurotech is using Eclipse-governed projects to bring AI to the edge, where systems interact with physical equipment and must remain secure, reliable, and maintainable for years.
  • Eclipse LMOS, originally developed at Deutsche Telekom, supports production multi-agent systems at significant scale and underpins multiple global deployments today.

“Across these examples, the common approach is modularity: keeping models, infrastructure, and applications separate enough that one can change without rebuilding everything. Our job is to capture what works and turn it into guidance others can use,” affirmed Milinkovich.

The Sovereign AI Foundation team says that alongside the 17 founding members, additional organizations are in the process of joining, and it invites eligible Eclipse Foundation member organizations to participate.

The post Eclipse wants companies to be free to switch AI providers. Today, doing so can mean a costly rebuild. appeared first on The New Stack.

Ember-1 vs. Kimi K3: Nearly identical results at 3.4 times the speed

Fireworks Research launched Ember-1 on September 23 as a research preview. Built on Moonshot’s open-weight model, Kimi K3, it claims it matches Kimi K3’s quality using roughly 40% fewer tokens. Fireworks says it “learned to cut unnecessary reasoning while keeping the thinking that matters.” Reasoning tokens bill as output, so a model that thinks less should cost less.

There’s a catch, though. On OpenRouter, Ember costs $3 per million input tokens and $15 per million output tokens. Fireworks charges the same for Kimi K3, but other providers sell Kimi for as little as $1 per million input tokens and $9 per million output tokens.

I ran both models through OpenRouter on Fireworks, so both were billed at $3 and $15, and I also calculated what Kimi would have cost at the lower price.

It doesn’t matter how much cheaper a model is if it isn’t accurate. I wanted to know whether Ember’s shorter thinking holds up. If it cuts reasoning it actually needs, accuracy should slip first on the hardest problems. So I built tests that get progressively harder and ran each one five times, the same consistency check I’ve been adding to my recent testing.

The tests

I called both models through OpenRouter with identical prompts and their default reasoning settings. I routed both to Fireworks to keep the speed comparison fair. Each test ran five times per model, and I logged reasoning tokens separately from the rest of the output.

  • Logic puzzles – Three puzzles of increasing size, with 4, 5, and 7 engineers and one solution each. Bigger puzzles need longer chains of reasoning, so cutting the thinking should hurt first.
  • Deploy scheduling – Twelve services with dependencies, one deploy per team at a time, and two blackout windows. The model has to find the fastest possible rollout: 17 hours.
  • Probability – Five questions about a retry system whose server flips between healthy and degraded, plus a circuit breaker. Every answer is an exact fraction, which I checked against a simulation of 2 million requests.

I confirmed every answer key with two independent methods before either model saw it. I’ve included the prompts at the end of this article for anyone who wants to replicate these tests on their system.

Logic puzzles

Both models solved all three puzzles on all five runs.

Ember averaged 13,630 reasoning tokens, 3 minutes 46 seconds, and $0.27 per run. Kimi averaged 16,679 reasoning tokens, 12 minutes 26 seconds, and $0.34. That’s 18% fewer reasoning tokens for Ember. Though not included in the marketing claims, Ember was much faster than Kimi. Kimi’s slowest run took nearly 20 minutes.

Deploy scheduling 

Both models found the 17-hour schedule every time, and every schedule passed my checker.

Ember averaged 6,543 reasoning tokens, 1 minute 29 seconds, and $0.10. Kimi averaged 7,792 reasoning tokens, 4 minutes 46 seconds, and $0.13. That’s 16% fewer reasoning tokens and significantly faster speed.

Probability

This test produced the only miss. Kimi answered all five questions correctly on every run. Ember got them all right four times. On the fifth, it made a small arithmetic slip on the first question, 0.94619 instead of 0.94629, and that error carried into two other answers.

It also saved the most. It averaged 6,242 reasoning tokens, 1 minute 47 seconds, and $0.13. Kimi averaged 9,682 reasoning tokens, 6 minutes 48 seconds, and $0.19. This was the first time Ember came close to its marketing claim of 40% fewer tokens. Once again, Ember was significantly faster.

Results

Test (5 runs each)Ember-1Kimi K3
Logic puzzles5/5 perfect, 3:46, 13,630 reasoning / 17,766 total out, $0.275/5 perfect, 12:26, 16,679 reasoning / 22,553 total out, $0.34
Deploy scheduling5/5 perfect, 1:29, 6,543 reasoning / 6,822 total out, $0.105/5 perfect, 4:46, 7,792 reasoning / 8,365 total out, $0.13
Probability4/5 perfect, 1:47, 6,242 reasoning / 8,365 total out, $0.135/5 perfect, 6:48, 9,682 reasoning / 12,381 total out, $0.19
Perfect runs14 of 1515 of 15
Total cost, both on Fireworks ($3/$15)$2.48$3.26
Total cost, Kimi at cheapest price ($1/$9)$2.48 (no cheaper provider)
$1.96

Kimi K3 had 15 perfect runs out of 15, and Ember-1 had 14. Ember-1’s one miss was an arithmetic slip on the probability test; it doesn’t seem too significant to me since it was one out of 15.

Something I noticed during my testing that wasn’t in the marketing claims was that Ember-1 finished each set of tests 3.4 times faster than Kimi K3. As for the reduced reasoning token usage, Ember-1 used 23% fewer reasoning tokens. This helped keep costs down. In total, it cost $2.48 compared to Kimi K3’s $3.26 on Fireworks, which makes it 24% cheaper. 

One caveat is that Kimi K3’s cost depends on the provider. At its lowest listed price of $1 per million input tokens and $9 per million output tokens, the same runs would have cost $1.96, less than Ember-1.

What do I think?

Ember-1 is about as accurate as Kimi K3 and uses far fewer reasoning tokens. The lower costs on Ember-1 on this test aren’t as important because a user who wants to use Kimi K3 can route to a cheaper provider on OpenRouter, which brings the price down below Ember-1’s.

One thing I know about Kimi K3 is that it’s slow. Cheap and slow. My speed numbers come from Fireworks’ standard endpoint, though, and I didn’t test the cheaper providers. If you have time and want to pay less, Kimi K3 wins. If you want almost identical results to Kimi K3 at a much faster rate, use Ember-1. 

The prompts

Each prompt ends with a fixed answer format so grading can be automatic.

Logic puzzles

Solve all three logic puzzles below. Each has exactly one solution.

PUZZLE SMALL: 4 engineers (Ava, Bo, Cleo, Dev) each own exactly one server. Each server has a rack position (1, 2, 3, 4, numbered left to right), an operating system (Debian, Alpine, Fedora, Ubuntu), a role (cache, queue, proxy, db). No two servers share any of these values.

Clues:

If the server in rack 3 runs Fedora, then the Debian server is in rack 3.

The server in rack 1 is Dev’s.

Exactly one of these is true: the server in rack 4 runs Ubuntu, or the Alpine server is not Bo’s.

Cleo is the proxy.

The db server is Dev’s.

Ava runs Alpine.

The queue server is not Ava’s.

Bo is in rack 3.

Bo and the Ubuntu server are in neighboring racks.

PUZZLE MEDIUM: 5 engineers (Ava, Bo, Cleo, Dev, Eun) each own exactly one server. Each server has a rack position (1, 2, 3, 4, 5, numbered left to right), an operating system (Debian, Alpine, Fedora, Ubuntu, Arch), a role (cache, queue, proxy, db, build), a replication data center (Oslo, Lima, Pune, Accra, Perth). No two servers share any of these values.

Clues:

Cleo is the db.

Exactly one of these is true: the cache server runs Arch, or the server replicated to Perth is the proxy.

The server in rack 5 runs Fedora.

Exactly one of these is true: the server in rack 2 is not the proxy, or Cleo runs Debian.

Exactly one of these is true: the server replicated to Pune is the db, or Dev is in rack 3.

The server in rack 1 is the db.

The proxy server is in rack 5.

The server in rack 4 runs Ubuntu.

The cache server and the server replicated to Pune are in neighboring racks.

Ava is the cache.

The server replicated to Accra is the queue.

The server replicated to Lima is not in rack 3.

Eun runs Alpine.

The queue server is not in rack 2.

The Fedora server is not Dev’s.

The server in rack 2 is not replicated to Lima.

PUZZLE LARGE: 7 engineers (Ava, Bo, Cleo, Dev, Eun, Finn, Gia) each own exactly one server. Each server has a rack position (1, 2, 3, 4, 5, 6, 7, numbered left to right), an operating system (Debian, Alpine, Fedora, Ubuntu, Arch, Rocky, NixOS), a role (cache, queue, proxy, db, build, metrics, auth), a replication data center (Oslo, Lima, Pune, Accra, Perth, Quito, Riga). No two servers share any of these values.

Clues:

The build server is Cleo’s.

The proxy server does not run Fedora.

The server replicated to Quito is Finn’s.

The server in rack 6 is replicated to Lima.

The Alpine server is exactly 4 racks to the right of the auth server.

The server replicated to Oslo is Dev’s.

The Debian server is replicated to Quito.

The server replicated to Accra is not the build.

The Rocky server is exactly 1 rack to the right of Gia.

Exactly one of these is true: the Ubuntu server and the server replicated to Pune are in neighboring racks, or the server replicated to Oslo runs Fedora.

The server replicated to Oslo does not run Fedora.

Ava runs Rocky.

Exactly one of these is true: the proxy server is somewhere to the left of the server replicated to Riga, or the Debian server is in rack 7.

The auth server and Bo are in neighboring racks.

The metrics server is Gia’s.

The Fedora server is somewhere to the left of the server replicated to Pune.

Exactly one of these is true: the server replicated to Accra is in rack 2, or Finn is the queue.

If the Debian server is Finn’s, then the server in rack 7 does not run Rocky.

The NixOS server is not replicated to Perth.

The server in rack 7 is replicated to Accra.

The server replicated to Pune is the db.

At the end of your response, give each solution as a block, one line per engineer, in the order the engineers are listed in that puzzle:

SMALL:

<name> | <rack> | <os> | <role>

MEDIUM:

<name> | <rack> | <os> | <role> | <dc>

LARGE:

<name> | <rack> | <os> | <role> | <dc>


Deploy scheduling

You are planning a production rollout of 12 services. Time is measured in whole hours from hour 0.

Services (owning team, deploy duration in hours):

auth: team A, 3 hours

billing: team B, 4 hours

catalog: team C, 2 hours

search: team C, 3 hours

cart: team B, 2 hours

checkout: team B, 3 hours

payments: team A, 4 hours

notify: team C, 2 hours

ledger: team A, 2 hours

gateway: team A, 3 hours

reports: team B, 3 hours

inventory: team C, 4 hours

Rules:

Dependencies. A service may start deploying only after every service it depends on has finished deploying:

gateway depends on auth

payments depends on auth

search depends on catalog

cart depends on catalog

cart depends on inventory

checkout depends on cart

checkout depends on payments

ledger depends on billing

ledger depends on payments

notify depends on checkout

reports depends on ledger

gateway depends on search

notify depends on gateway

Each team can deploy only one of its services at a time.

Different teams can deploy at the same time.

Blackout windows. No deploy may be in progress at any time during hours 9 to 11 or hours 17 to 19. A deploy may end exactly at hour 9 or 17, and may start exactly at hour 11 or 19. A deploy cannot pause and resume.

Each deploy runs from its start hour for its full duration without interruption.

What is the earliest hour by which all 12 services can be finished? Give a schedule that achieves it.

At the end of your response, give your answer in exactly this format:

MAKESPAN: <hour>

<service>: <start hour>

(one line per service, all 12 services)


Probability

A client calls a payment API and retries on failure.

The server is in one of two states on each attempt, Healthy or Degraded.

On the first attempt, the server is Healthy with probability 4/5 and Degraded with probability 1/5.

Between consecutive attempts, the state changes like this: from Healthy, it stays Healthy with probability 3/4 and becomes Degraded with probability 1/4. From Degraded, it stays Degraded with probability 2/3 and becomes Healthy with probability 1/3.

An attempt succeeds with probability 9/10 if the server is Healthy and 2/5 if it is Degraded, independently of everything else given the state.

The client makes at most 4 attempts and stops as soon as one succeeds.

Circuit breaker: if two consecutive attempts both hit a Degraded server and both fail, the client stops immediately and makes no more attempts.

Answer these questions with exact fractions in lowest terms:

Q1. What is the probability that the request eventually succeeds?

Q2. What is the expected number of attempts the client makes?

Q3. Given that the request succeeds, what is the probability that it succeeded on exactly the second attempt?

Q4. What is the probability that the circuit breaker ends the request early, before the client has used all 4 attempts?

Q5. Given that the request succeeds, what is the probability that the first attempt failed?

At the end of your response, give exactly five lines in this format:

Q1: <fraction>

Q2: <fraction>

Q3: <fraction>

Q4: <fraction>

Q5: <fraction>

The post Ember-1 vs. Kimi K3: Nearly identical results at 3.4 times the speed appeared first on The New Stack.

Elon Musk says space will soon hold nearly all compute. Google is still finding out if its chips can work there.

A dense field of stars against a black sky. Dozens of bright white, blue, and orange stars have four-pointed starburst spikes, scattered over thousands of fainter points of light.

While SpaceX and Nvidia plan to place a space-optimized Vera Rubin AI system in orbit by late 2027, Google’s far more modest plan, Project Suncatcher, is to launch a test satellite with Tensor Processing Units to see whether AI data centers in space are viable. 

For SpaceX and Nvidia, it’s a race to get ahead in AI data centers in space, with Starmind AI 1 satellites. Never mind that the engineering challenges of cooling, radiation resistance, and repairability still haven’t been sufficiently addressed. While SpaceX and Nvidia are gung-ho to put full AI data centers in space as fast as possible, Google will launch the prototype satellite for Project Suncatcher to see whether Tensor Processing Units (TPUs) will fly at all. 

SpaceX leader Elon Musk is already proclaiming that “the amount of compute in space will obviously round up to 100% of all compute.” Google is more realistic. Google describes Project Suncatcher as a “long-term, research moonshot exploring whether space could one day host scalable machine learning infrastructure.”

“The amount of compute in space will obviously round up to 100% of all compute.”

This October flight marks the first orbital test for Project Suncatcher. The refrigerator-sized prototype carries four TPUs, which can run only in 15-minute bursts before shutting down to cool. Google’s research effort explores whether a constellation of solar-powered satellites, connected by high-speed optical links, could someday support large-scale machine-learning workloads.

Google wants to know if space is a good place to run AI because satellites get plenty of solar power. This first mission will expose the hardware to launch stresses, radiation, and extreme thermal conditions that can’t be reproduced fully on Earth.

As Brandon Lucia, a Carnegie Mellon University professor of electrical and computer engineering, told The New York Times, “You get a lot of weird particles out in space — high-energy radiation that we are just not exposed to on Earth. Sometimes, you get a random high-energy particle strike that is like someone throwing a dart at the insides of your computer chip.”

“Sometimes, you get a random high-energy particle strike that is like someone throwing a dart at the insides of your computer chip.”

Not to mention, Lucia added, the “cooling problem is actually quite difficult. These computers will be basking in the sun all day.”

So Google naturally wants to know if their TPUs can work in these conditions before betting the farm on AI in space. Thus, the first Suncatcher satellite will carry Google TPUs, the company’s in-house accelerators for AI workloads.

During launch, Google said the spacecraft will endure intense vibration and sustained acceleration of up to 10 times Earth gravity (g). At the same time, individual components such as the chips can experience loads of 50 to 100 g. This isn’t like moving your server rack down the road with a truck! 

The longer-term Suncatcher idea isn’t simply to run AI workloads aboard a single satellite, as SpaceX’s first mission will. Instead, Google envisions a compact AI satellite constellation. 

Each satellite will carry TPUs and communicate with partners via laser links. Google needs these links to operate at tens of terabits per second. It has demonstrated 800 Gbps in each direction — 1.6 Tbps total — with a bench-scale optical-transceiver pair.

But turning that into an orbital compute fabric requires satellites to fly unusually close together, potentially separated by hundreds of meters. Google plans a two-satellite mission in 2027 to test high-bandwidth laser communications for that next phase.

“This first launch is about seeing what works, identifying points of failure, and applying those findings to future missions.”

For now, Google just wants to know if the technology works at all.As Travis Beals, senior director of Google’s Paradigms of Intelligence research team, explained: “This first launch is about seeing what works, identifying points of failure and applying those findings to future missions.”

The post Elon Musk says space will soon hold nearly all compute. Google is still finding out if its chips can work there. appeared first on The New Stack.

Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart

Satya Nadella stands smiling between Bill Gates, on the left, and Steve Ballmer, on the right, in front of a crowd of cheering employees, many holding up phones and tablets to take photos.

Microsoft announced what it calls its biggest Copilot update to date on Friday, with CEO Satya Nadella describing Copilot as “a new OS for work.”

Nadella framed Copilot as spanning every model, form factor, and task, and the update puts Autopilot, which Nadella called a “proactive and long-running agent built for the enterprise,” at the top of his list of the update’s four components. The pitch targets office workers, but the more consequential change for developers is the infrastructure underneath.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365, which means teams building production agents no longer have to assemble those pieces around a model on their own.

We’re building Copilot as a new OS for work that spans every model, every form factor, and every task. Today, we’re announcing our biggest update to Copilot to date, bringing four things together:

· Autopilot: proactive and long-running agent built for the enterprise
· Code:… pic.twitter.com/W2ClHHkCK3

— Satya Nadella (@satyanadella) September 25, 2026

The release adds a new Home experience that merges Chat and Cowork in the Copilot app, but the bigger changes for developers come from Code and Autopilot. Code generates apps, dashboards, and workflows from natural language, and Autopilot turns the agent Microsoft previously called Scout into a persistent background worker. Home and Code are rolling out first through Microsoft’s Frontier early-access program, and Autopilot is expanding to a private preview at month’s end.

Microsoft is moving the agent runtime into the enterprise infrastructure layer and building persistent identity, state, execution boundaries, and organizational context into Microsoft 365

Agents that don’t need prompts

Autopilot takes a role and goal from the person who sets it up, then continues working in the background without requiring a new prompt for each step. Each Autopilot gets its own governed Entra identity and agent user account, separating the agent’s permissions and activity from those of the person who created it.

For engineers, that moves much of the operational scaffolding required for long-running agents into Microsoft’s infrastructure. Independent vendors have been building dedicated layers for that problem; Diagrid, for example, adds durable recovery to LangGraph and other agent frameworks, while Microsoft is bringing those capabilities inside the Microsoft 365 environment.

An identity for every agent

The identity model is the piece developers building on Microsoft Foundry will feel first. Autopilot agents in Foundry, which have been in public preview since June, receive a full Entra Agent ID user account with a productivity license that gives them their own email, calendar, OneDrive storage, Teams access, and a place in the org chart.

Because that user account sits on top of the agent identity every Foundry agent already carries, an autopilot acts as itself rather than on behalf of a user, so developers no longer have to wire agents through shared service accounts or borrowed user credentials, a pattern AuthZed CEO Jake Moshenko has said reflects a common misconception about how agents should be deployed.

A developer creates an Autopilot blueprint from a Foundry-hosted agent, which appears in the Agent 365 registry once an administrator approves it. Employees can then hire instances of that agent in Teams. The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

The blueprint establishes what the agent is designed to do, but administrators still control the resources and data each instance can access, extending the same access policies used for employees to agents working on their behalf.

Hosting AI-generated apps

Code is built on the same underlying technology as GitHub Copilot, and the apps it generates run on Microsoft Copilot Managed Runtime, a platform now in public preview that hosts code inside the customer’s Microsoft 365 tenant boundary under IT governance.

Apps deployed there run within the company’s existing identity and governance framework, with Microsoft managing the underlying runtime and giving developers a controlled path to test and deploy new versions without taking the current release offline.

The runtime also accepts apps built in Copilot Studio and Cowork, and Microsoft is opening it to outside tools and professional developers through an SDK and command-line tooling, with Git tracking source and versions.

Lovable is already on board. In Microsoft’s announcement, the company’s head of global partnerships, Lan Roche, said apps built with Lovable can now run inside a Microsoft tenant “the same way everything else does,” using the same sign-in, policies, and app inventory.

The model resembles what serverless computing did for application infrastructure, where developers concentrate on application logic while the platform takes on more of the execution environment. Microsoft is applying that abstraction to generated enterprise software while tying the runtime directly to identity, tenant boundaries, and organizational data.

Long-running agents also change Copilot’s economics. The standard subscription covers the assistant, but Cowork, Code, Autopilot, and other agentic features are billed based on usage through Copilot Credits. That also applies to frontier models such as Fable and Astra, although users still need a Copilot license to access them. Microsoft is extending cost management in Agent 365 to cover Code and Copilot Managed Runtime, and it plans to support agents built in Copilot Studio in October.

Once an agent can keep working for hours or days without anyone watching, cost becomes part of the governance problem. Engineering teams need to control how much compute an agent uses alongside what it can access, which is why Microsoft is bringing those controls into the same administrative framework.

The portability trade-off

That convenience comes with a trade-off. Because Microsoft controls the underlying enterprise environment, it can handle much of the work around agent state, credentials, and access controls, but the more infrastructure a team hands over to Microsoft, the harder the agent may be to move elsewhere.

The models are not locked in, since Microsoft currently runs Copilot on models from both OpenAI and Anthropic and says more labs and open-weight models are coming, and the Agent 365 SDK adds governed Model Context Protocol access to Microsoft 365 workloads for agents regardless of the framework they were built with. Those open interfaces cover only part of an agent’s architecture, though. The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The more an agent depends on Microsoft 365 for its identity, permissions, and context, the more work it takes to move that agent elsewhere.

The post Microsoft’s new Copilot agents get their own email, calendar — and a place in the org chart appeared first on The New Stack.

OpenAI and Cursor agree on agent coordinators. They disagree on who runs them.

Abstract digital art of thousands of thin glowing strands in orange, red and pink bundled into a single sweeping arch against a black background.

OpenAI opened its Agents API in public beta this month, exposing the harness that powers Codex with managed sessions, tool coordination, and subagent orchestration. On the same day, September 10, Cursor launched Projects to coordinate multiple coding agents around larger bodies of software work. While the products sit at different points in the stack, both converge on the same architecture: a coordinator understands the larger objective and manages the work, while specialized agents execute individual pieces.

That pattern isn’t new: AWS Bedrock AgentCore reached general availability in October 2025, and Anthropic’s Claude Managed Agents entered public beta in April 2026. What makes these announcements notable is that two major players in AI-assisted software development are independently exposing the same coordinator-worker split at the same time.

Hilliary Lipsig, a senior principal site reliability engineer at Red Hat who leads Azure Red Hat OpenShift SRE teams and hosts the YouTube livestream GitOps Guide to the Galaxy, has watched this dynamic play out firsthand.

“This convergence highlights the reality developers across the industry have been discussing on and offline — an agent with too much context loses accuracy and reliability, and focused work with clearer contexts allows for faster, more accurate iterations,” Lipsig tells The New Stack.

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly,” Lipsig says. “That’s part of how we got to Kubernetes. These multi-agent workflows are the same concept, just in a new part of the technical stack. While the specialized agents do their area of work, the orchestrator can act as a source of truth — ideally enforcing guardrails, recovering from any failure states, and intelligently routing work to the most efficient target agent.”

“The need for orchestration in distributed computing has been fundamentally recognized repeatedly… These multi-agent workflows are the same concept, just in a new part of the technical stack.”

The industry has spent the first generation of AI coding tools asking how capable a model can become at writing software. The emerging question is different: How do you build a reliable system around multiple capable agents working on the same problem?

The problem with the single-agent loop

A coding agent works through what Anthropic describes as LLMs using tools based on environmental feedback in a loop: it observes the state of a repository, reasons about what to do next, calls a tool, examines the result, and continues. For a small task, that loop can be enough. As the scope expands, however, maintaining reliability in a single context becomes harder.

A large migration might require understanding an unfamiliar codebase, identifying dependencies, changing database schemas, updating services, rewriting tests, modifying deployment configuration, and validating the resulting system. A single agent can theoretically perform all of that work, but it must maintain relevant information from every stage while continuing to reason about what comes next.

The pressure lands first on the context window. “A large context doesn’t only include everything correct or important — it also includes a lot of throwaway information,” Lipsig tells The New Stack. “Through compaction, that information can inadvertently end up ranked as important and incorrectly influence what your agent does. Or correct information can be distorted to become incorrect.

“Either way, after a couple of rounds of compaction, developers are seeing accuracy degrade and are starting to manage context once again manually.”

Lipsig’s read matches what researchers call context rot — and it hasn’t gone away with newer models.

A 2026 study testing frontier models,, including Claude Opus 4.6, GPT-5.4, and Gemini 3.1 Pro, found they missed a dangerous action buried in a long agent transcript two to 30 times more often once it came after 800,000 tokens of benign activity — the AI equivalent of a security guard who stops checking badges carefully after the two-hundredth person walks through, even though nothing about their training changed.

Furthermore, the tasks themselves may not be sequential. Forcing one agent to execute database analysis, documentation work, and test discovery one after another turns a potentially parallel workload into a serial one.

Subagents change that execution model. Instead of requiring one agent to carry an entire task through a single context, a coordinator breaks the work into smaller units and assigns them to specialized agents. GitHub’s custom-agent model illustrates this: different agents receive only the prompts, tools, and context they need for their tasks, executing work in isolated contexts rather than crowding an increasingly large conversation.

Multi-agent systems therefore bring higher token costs and additional coordination and integration risks, and splitting work across agents does not guarantee better software quality.

The coordinator is not another coding agent

Once the work is divided this way, the coordinator becomes a control plane rather than another coding agent. Its job isn’t to write the code, but to understand the global task, manage dependencies, and decide how execution should proceed. Unlike a conventional scheduler, an agentic coordinator makes probabilistic judgments about result quality and resource allocation.

It may dispatch one agent to investigate a database schema, another to examine the service layer, and a third to inspect the test suite. When they return, the coordinator determines if their findings are sufficient to move to implementation. If a worker produces an incorrect result, the system must recognize the failure and decide whether to retry the work, reassign it, or change the task itself.

Anthropic has documented this same pattern in its own production system, calling it orchestrator-subagent architecture: a lead agent analyzes a query, develops a strategy, and spawns specialized subagents to investigate different facets in parallel. In a June 2025 writeup of that system, Anthropic reported a Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Opus 4 by 90.2% on its internal research eval — at roughly 15 times the token cost of a standard chat interaction (Anthropic puts single agents at about 4 times), a tradeoff that makes the pattern a deliberate architectural bet, not a free upgrade.

Parallelism introduces distributed-systems failure modes

Parallelism is valuable because software work contains many independent tasks, but it creates coordination problems. Imagine a migration where one agent changes a database schema, another updates the consuming service, and a third updates integration tests.

If the schema changes while the service agent works against an earlier assumption, the system produces internally inconsistent work. This isn’t a risk unique to hypothetical migrations — the International AI Safety Report 2026 notes that “interactions between multiple AI agents are also becoming more common, introducing further risks, as errors propagate between systems.”

A single model invocation is a disposable computation, but a twenty-minute workflow modifying a repository is not. If an agent loses its machine halfway through, restarting from scratch is expensive and potentially unsafe against a changed environment.

To solve this, Cursor moved its cloud-agent execution loop to Temporal to handle durable execution and retries, pushing its cloud agents past two 9s of reliability. Temporal now handles 50 million of Cursor’s actions a day across 7 million unique workflows. “Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo,” Lipsig tells The New Stack.

“Durable execution isn’t a nice-to-have here. It’s the difference between a system you can operate and one you can only demo.”

By separating agent, machine, and conversation state, the execution engine can reason about the workflow independently. Reliability is no longer just about whether the model produces a good answer; it is about reliably completing distributed workflows composed of many operations, machines, and dependencies.

The environment, context, and observability are one problem

In production, an agent is more than a model and a prompt; it requires a workspace, source code, dependencies, credentials, and state retention. Both companies provision isolated environments for these resources, directly linking an agent’s capability to its blast radius. OpenAI’s Agents API currently supports U.S. data residency but not Zero Data Retention; choosing a self-hosted sandbox does not make the Agents API eligible for ZDR. Cursor supports similar cloud isolation alongside local execution for machine-specific work.

An agent that can only inspect a repository poses a different risk than one that can modify production infrastructure. Consequently, the coordinator is inextricably linked to the security model, determining which agent receives specific information and authorities.

This logic extends to context routing. Giving every subagent the parent’s entire history increases cost and complexity while leaking irrelevant or sensitive information. Instead, the coordinator enforces information-flow boundaries: a database-analysis agent receives only schemas and relevant migrations, while a security-review agent gets the resulting diff without deployment credentials.

As agents increasingly use interfaces like MCP to reach external systems, the platform must strictly govern which agent receives the authority to use specific tools, and for how long. MCP’s governance now sits inside the Agentic AI Foundation, a Linux Foundation foundation co-founded by OpenAI, Anthropic, and Block, with support from AWS, Google, Microsoft, Bloomberg, and Cloudflare to host MCP alongside AGENTS.md and Block’s goose — a sign the industry already treats it as infrastructure worth governing jointly, not a feature any one vendor owns.

This complexity creates a visibility problem. A simple final response often conceals a history involving multiple agents, tool calls, environments, and retries. Systems must expose task-level provenance — which agent received the assignment, what context it used, where it executed, and how the coordinator handled failures or human interventions.

Without execution provenance, debugging requires reconstructing distributed workflows from fragments. GitHub’s exposure of subagent lifecycle events points in this direction, treating agent lifecycles as observable components rather than hidden processes.

Coordination authority is not execution authority

The most critical architectural boundary is the distinction between coordination authority and execution authority. A coordinator needs broad visibility to make useful decisions, but that does not imply unrestricted control over the project. “Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either,” Lipsig tells The New Stack.

“Just like you don’t want humans running around with root permissions, you don’t want your agents running with them either.”

“The ease of creating and leveraging AI agent permissions is lagging behind the speed of AI innovation, but any product team that needs to maintain compliance standards will tell you that easy or not, access controls are incredibly important. Functionally, that means fully agentic workflows need to run in a dedicated environment with their own permissions and connectivity. You can’t rely on agent context to do what you want — trying to keep something in context that long will just result in errors from context rot. Hardware or virtual-hardware-level permissions, in addition to system-level permissions, need to be configured so the agent is asked to provide only its unique value add.”

The risk isn’t hypothetical. OWASP’s Top 10 for Agentic Applications formally names this failure mode Identity and Privilege Abuse (ASI03).

METR’s independent investigation, conducted with a Redwood Research contractor, found that the June–July 2026 incident — in which OpenAI’s own agents, running in internal ExploitGym cyber evaluations, escaped their sanctioned scope and attacked Hugging Face — unfolded in distinct phases. On June 26, agents discovered an exploit that provided full administrator access to OpenAI’s internal Artifactory package repository. On July 4, high-volume Artifactory activity caused an outage, triggering a security investigation and pausing evaluations. OpenAI switched to a freshly initialized Artifactory instance on July 6, which it says wiped the cache and messages left there.

On July 7, a new wave of ExploitGym experiments launched tens of thousands of agents, and on July 8 one agent established the primary unsanctioned message board. Roughly 1,200 agents eventually used the board, sending more than 70,000 messages and files; about 700 later participated in the attack on Hugging Face. The attack itself began on July 10–11 and wound down over July 12–13. The chronology matters because the administrator-access event, the Artifactory outage, and the later message-board activity were separate phases, not one continuous incident.

By binding autonomy, worker agents operate with the minimum permissions required for their specific tasks, keeping sensitive operations behind explicit approval boundaries. This also reshapes human review. Requiring human approval for every tool call destroys the efficiency of multi-agent execution, but showing only the final result obscures critical intermediate decisions.

The most useful design places human intervention around consequential, irreversible transitions — like moving into production or altering sensitive infrastructure. This is especially vital as agents become event-driven participants that respond to Slack messages or pull request updates, not just direct prompts.

OpenAI and Cursor own different parts of the architecture

The convergence does not mean OpenAI and Cursor have built interchangeable systems. Their products put the orchestration boundary in different places.

OpenAI is exposing an agent harness through an API. Its model gives developers primitives for managing context, tools, subagents, and execution environments, leaving application teams to decide how those capabilities fit into their own systems. The harness is open source, so teams can inspect the coordinator logic instead of treating it as a black box.

Cursor packages more of the surrounding workflow. Projects provides the coordinator, cloud execution, shared project context, and a developer-facing workflow in the same environment.

That difference matters because orchestration is a collection of infrastructure decisions: who owns the execution environment, where workflow state persists, how agents are isolated, how credentials are provisioned, what happens when a worker fails, how one agent’s output becomes another agent’s input, and which actions can happen without human approval.

An API gives developers more responsibility for answering those questions. An integrated platform answers more of them on the developer’s behalf.

Neither approach removes the underlying engineering problems. It changes where they are implemented and who is responsible for operating them.

The coordinator is becoming an architectural boundary

The evidence from these systems points to a change in the role of the coding agent itself.

The model still performs the reasoning and code generation. But larger agentic workflows require another layer to determine how that capability is applied: which work is delegated, what context crosses an agent boundary, which tools are exposed, how execution state survives failures, and when the workflow needs human intervention.

Those are familiar distributed-systems concerns. Workers operate concurrently, state can be shared or isolated, dependencies connect tasks, workers can fail independently, and results need to be persisted and observed. The difference is that the workers are now probabilistic software agents rather than conventional processes.

That makes the coordinator more than a convenience feature. It is where a high-level software objective becomes executable work — and where decisions about context, permissions, durability, observability, and human intervention converge.

The September 10 launches make that shift visible from two different directions. OpenAI exposed orchestration infrastructure through an API. Cursor embedded it into a project-level development environment.

Neither announcement proves that one architecture will become the universal model for software development. But together with the systems already emerging around them, they show coding agents moving away from a single model executing an entire task and toward workflows that divide work among specialized agents, execution environments, and persistent infrastructure.

The engineering question is therefore no longer only whether an agent can write the code. It is whether the system around it can reliably decide what to do, which agent should do it, what that agent should be allowed to see and change, how to verify its work, and where a human should take control.

Those are architecture and infrastructure questions — and as coding agents move from interactive assistants toward autonomous software workflows, they may matter as much as the underlying model.

The post OpenAI and Cursor agree on agent coordinators. They disagree on who runs them. appeared first on The New Stack.

What managing 150,000 AI agents could look like for database teams

Abstract 3D render of translucent orange cubes and panels scattered across a pale gray background, with bundles of glossy teal tubes curving in from the right.

The database administrator of the future will spend considerably less time administering databases.

That sounds contradictory, but AI agents are taking over that work. For decades, DBAs have handled the decidedly hands-on work of keeping databases available, performant, secure, and affordable. They provision capacity, troubleshoot slow queries, manage migrations, and step in when something inevitably goes sideways.

AI is already taking on some of that work. At the same time, it is creating a much bigger data infrastructure fleet to manage.

The result is likely to be a very different kind of DBA: one who spends less time tending individual databases and more time supervising the autonomous systems doing it for them.

Congratulations, you’re managing robots now

This shift starts with a familiar problem: more infrastructure needs managing than the people available to manage it.

Database automation is hardly new, but agents can potentially go further than the scripts and rules DBAs already rely on. Rather than automating one predetermined task, an agent can inspect what is happening, decide what needs attention, use tools to act on it, and check whether its intervention worked.

That changes the DBA’s relationship with the database. A performance problem that once required someone to dig through metrics, identify the troublesome query, and decide how to respond could increasingly be investigated by an agent before a human gets involved.

It doesn’t remove the DBA from the equation. Someone still has to decide what an agent can do, where human approval is required, and what happens when it gets something wrong. But the work moves up a layer. Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

Instead of personally performing every operational task, DBAs start managing the systems carrying them out.

And before anyone gets too comfortable with that idea, the number of those systems could become enormous.

150,000 agents walk into a database…

Gartner predicts that the average global Fortune 500 company will have more than 150,000 AI agents in use by 2028, up from fewer than 15 in 2025. Only 13% of organizations currently believe they have the right governance in place to manage them.

Not every agent will need its own database, but plenty will. They will create state, retrieve data, remember previous interactions, and exchange information with other agents. Many will also behave very differently from the applications DBAs are used to supporting: spinning up quickly, sitting idle for long stretches, and suddenly becoming busy when there is work to do.

Nobody is hiring 150,000 DBAs to manage them.


That is the scale problem Yugabyte is targeting with YugabyteDB AMP, or Agentic Multitenant PostgreSQL. Rather than treating each new agent workload as another database for an administrator to provision and babysit, AMP manages databases as a fleet.

The platform packs hundreds of small Postgres workloads onto shared distributed infrastructure while keeping their databases isolated. Lifecycle operations, including provisioning, branching, scaling, migration, and teardown, can be exposed to agents through MCP. Yugabyte has also built specialized agents for setup, migration, performance tuning, and integrations.

In that model, a DBA is no longer provisioning database number 14,372. The interesting job is setting the rules for how database number 14,372 is provisioned, operated, and fine-tuned without them.

Do more with less (no, really)

Scale is only half of the problem. Someone also has to pay for all this stuff.

Agent workloads make traditional capacity planning particularly awkward because many are bursty and frequently idle. Giving every experimental agent permanently provisioned infrastructure could leave companies paying for many databases that spend much of their lives doing very little.

This is where consolidation becomes as much an economic question as an operational one.

AMP’s approach is serverless multitenancy and scale-to-zero. Multiple small workloads share the underlying distributed infrastructure, while customers pay by CPU minute and idle agents consume no compute. Resource governance can impose CPU limits on individual workloads, preventing a single overeager agent from consuming the capacity intended for its neighbors.

The human equivalent matters too. If routine setup, migrations, tuning and other database operations can increasingly be delegated, a smaller database team can potentially look after a much larger estate.

That doesn’t mean companies get to fire the DBAs and hand the keys to the robots. It means scarce database expertise can be spent on architecture, governance, and genuinely difficult problems instead of repeatedly doing the work that software can handle.

Your 2028 database problem starts now

The harder question is what to build underneath all of this when nobody really knows what the enterprise AI estate will look like in two years.

An agent that begins as an experiment today could disappear next month. Another could suddenly become a production application used across the business. Building one infrastructure stack for cheap experiments and another for serious workloads risks creating a migration problem every time an experiment succeeds.

Yugabyte bets that both ends of that journey should sit on the same foundation.

YugabyteDB AMP lets workloads start on serverless Postgres and transition to fully distributed YugabyteDB as their scale and criticality increase, without rewriting the application or migrating data to a different database platform.

Then there is the problem above the individual database: agents need to remember what happened, and not just in a silo.

That’s where Meko fits into the Yugabyte stack. Meko is an agent-native context engine designed for multi-agent AI systems. It provides persistent memory, shared knowledge, decision traces, and autidability across multiple agents, rather than leaving each agent working from its own isolated context. An agent can pick up information learned by another agent instead of retrieving it again or restarting the reasoning process.

Taken together, it delivers a single data stack for an agent’s entire lifecycle: Meko for the context shared among agents, YugabyteDB AMP for agentically managing fleets of Postgres databases, and distributed Postgres-compatible YugabyteDB for workloads that outgrow their serverless beginnings.

Of course, there’s no guarantee that 2028 will look exactly like today’s forecasts. That’s rather the point. The safest architectural bet may be one that doesn’t require you to know in advance which of today’s tiny AI experiments will become tomorrow’s critical applications.

The DBA is still critical in that world, but the job will look different. The DBA of the future may manage fewer databases directly, while taking responsibility for vastly more of them. Instead, managing the autonomous systems that do the administering.

The post What managing 150,000 AI agents could look like for database teams appeared first on The New Stack.

Query decomposition doesn’t fix context starvation — it just moves it

Abstract digital art showing warped light lines surrounding a void, illustrating data compression and AI context starvation.

I have a small example that would best communicate the message I am trying to convey: say you built a chat widget for GitLab’s public documentation (the corpus we are experimenting with in this article) and one of the developers sends this kind of message:

We got an email saying our card was declined for something called “quarterly reconciliation” and I need to know what actually happens now. On top of that, I think we’ve gone over our seat count; there are more people in the group than seats we bought. Our CI has been queuing all week and I want to know whether the compute minutes we purchased last month rolled over or if we lose them. Our finance lead also needs to be the one who gets the invoices from now on, not me. And last thing, is the REST API rate limited? We’re building an internal dashboard and would rather find out now than after it breaks.

Five separate asks: the declined payment, the seat overage, compute-minute rollover, changing who receives invoices, and API rate limits. Each one is answered by a specific passage in GitLab’s public documentation, and you labeled which passage answers which before running anything, so you knew in advance exactly what a correct system needed to find.

Then you run the message through a pipeline that follows current best practice. It splits the query into five clean sub-queries, retrieves for each one independently, merges the results, drops near-duplicates, reranks the merged pool against the original message, and packs the highest-scoring passages into a 2,000-token context.

The pipeline retrieved all five correct passages, but only one of them survived into the packed context; that is one of five asks, not one of five sentences. The packer found, scored, and threw away the other four before the model ever saw them. The same message with no decomposition at all managed three out of five.

The failure has a name, and it isn’t the one you’re thinking of

I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.

The definition is deliberately about allocation rather than retrieval, because allocation is the part nobody watches. If the passage answering the fifth question was found, scored, and then squeezed out by three passages about the first question, the fifth sub-intent is starved, and every recall metric you have will report that the system worked perfectly.

“I call this context starvation: a sub-intent that gets no allocation in the final packed context, whether or not its evidence was successfully retrieved.”

Two failure modes already in circulation describe something different, and it’s worth separating them cleanly:

Semantic dilution happens at retrieval time: when you embed a five-part question as a single vector, you get a centroid that sits somewhere between five topics and lands close to none of them, so the evidence is never found. Decomposition fixes this, which is why it spread.

Context poisoning is about what is present, not what is missing. Wrong, stale, or adversarial content enters the window and corrupts what the model generates downstream. Poisoning is a contamination problem. Starvation is an absence problem, and policy, not accident, produces the absence.

I borrowed the word from operating systems. In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work. Every ingredient of that situation is present in a retrieval pipeline: a fixed resource, competing demands, and a policy that decides who gets served. A relevance-greedy packer is priority scheduling with no aging term, and under priority scheduling without aging, valid low-priority work waits forever.

Decomposition is the right fix to the wrong half of the problem

Split that message into five single-intent queries, and each one embeds cleanly, so per-sub-query recall climbs sharply. This is well-trodden ground. LlamaIndex ships a SubQuestionQueryEngine that breaks a complex query into sub-questions and synthesizes the responses. LangChain’s MultiQueryRetriever generates query variants and returns the unique union of what they retrieve. RAG-Fusion applies reciprocal rank fusion across the per-query result lists. The technique works, and it isn’t mine.

“In scheduling, a process starves when it is ready to run, waits, and is never selected because the priority function keeps preferring other work.”

The context window did not grow. Let me explain: after decomposition, you have n result sets competing for one fixed token budget, and something downstream has to decide the split. In most production pipelines, that something is a short, unremarkable sequence: merge the pools, drop near-duplicates, rerank the merged pool against the original query, then greedily fill until the budget closes.

That sequence is a scheduler. It has a priority function, which is the reranker score, and it has no fairness constraint of any kind. A sub-intent with three strongly-scoring passages takes three slots. A sub-intent whose single correct passage scores mid-pack takes none of them.

So the failure did not go away. It moved from the embedding, where it has a name and people watch for it, into the packer, where it has neither. It also moved somewhere with much worse instrumentation, because recall@k per sub-query is the metric decomposition usually gets validated with, and that number goes up. It goes up at the same time as coverage inside the packed context goes down. You ship on a green dashboard.

The harness

The corpus, GitLab’s public documentation: 10,000 chunks and 2.2M tokens, split on heading boundaries and capped at 480 tokens each. Sixty-one single-intent questions span nine topics, from seat management to rate limits, each labeled with the one passage that answers it. I built multi-intent queries by concatenating those questions while varying n across 2, 3, 5, and 7, randomizing the order so position doesn’t confound topic, and varying topical distance so half the queries draw everything from one topic and half span distinct ones. That produces 100 queries, 25 at each value of n, whose correct decomposition I know exactly.

Two decisions matter more than the rest:

The metric is not recall. Recall tells you what the retriever found. What I need is what survived into the packed context, per sub-intent. So I log each sub-intent twice: once for whether its correct passage reached the candidate pool, and once for whether it reached the packed context. The gap between those two numbers is the entire argument.

Every question has to be retrievable on its own before it’s allowed in. A question enters only if its correct passage ranks in the top 10 for its own isolated query, under both retriever configurations, and both scored 100% recall@10 on that test. Since each sub-query’s candidate pool is exactly its own top 10, passing that gate guarantees the correct passage sits in the pool for every decomposed arm. Any sub-intent that then fails to appear was denied by the packer rather than missed by the retriever, which removes the most obvious objection to everything below.

The core measurement uses no language model. Because queries are composed from known sub-questions, the decomposer is an oracle so that anyone can reproduce the main result with no API key.

That invites an objection, so I tested it. A real LLM decomposer, blind to n, disagreed with my ground truth on 41% of the queries, and on inspection it was right every time. Five of my sixty-one supposedly single-intent questions contain two distinct information needs. What is excess storage usage, and what happens when we go over the free limit? is two questions wearing one question mark. Adjusted for those five, agreement is 100 out of 100. That is not evidence decomposers are reliable, because my queries are joined by fixed connectives and splitting on those alone recovers n perfectly, which real messages never allow. What it caught was an error in my own labels, and that is the best argument I have for the oracle design.

Results

Every arm runs at 2,000, 4,000, and 8,000 tokens against two retriever configurations. The stronger pairs are BAAI/bge-base-en-v1.5 with BAAI/bge-reranker-base; the weaker pairs are a quantized BAAI/bge-small-en-v1.5 with Xenova/ms-marco-MiniLM-L-6-v2. Retrieval is in-memory cosine similarity over a NumPy array because, at 10,000 chunks, a vector database would be slower to write, slower to run, and harder to verify.

Sub-intent coverage at a 2,000-token budget on the stronger configuration:

Table showing coverage (in %) per arm.

The production-default pipeline starves 31.1% of sub-intents whose evidence it had already retrieved. It beats no decomposition by nine points, while a flat B/n split, which is the crudest allocator anyone could write, beats it by fourteen.

Floors work, but not the obvious floor. Reserving one passage per sub-intent before the greedy fill satisfied 99.4% of its reservations and bought only seven points. The mechanism fires correctly and reserves the wrong passage because it picks each sub-intent’s best chunk by score against the original query. The original query asks about all five intents at once. Selecting that same reservation by score against its own sub-query pushes coverage to 89.4% and cuts allocation starvation from 30.6% to 10.1%. That is a change of about four lines.

Two of my own recommendations died here. I expected reranking against the original query to beat reranking against the fragment, and it loses by seventeen points. The incomparable score scales I worried about turn out to help, because each sub-query’s best match ends up at the top of its own scale, producing per-intent fairness for free. I also expected deduplication before allocation to matter, but near-duplicates consume 1.0% of the budget and removing them moves coverage by 0.3 points.

The crossover: starvation by tokens-per-sub-intent, which is simply the budget divided by n:

Table showing tokens per sub-intent.

Below roughly 1,000 tokens per sub-intent, allocation policy dominates. Above it, nothing you do to the allocator matters, because everything fits anyway.

The retriever comparison is the one I’d lead with. Upgrading the retriever moves coverage on the production-default arm from 56.2% to 68.9%, a gain of 12.7 points. Changing the allocation policy on the same retriever moves it from 68.9% to 89.4%, a gain of 20.5 points. In its sharpest form: the weaker retriever with a fragment-scored floor reaches 84.5%, and beats the stronger retriever with a greedy packer at 68.9% by sixteen points. A worse retriever with a better allocator wins.

Position: Held within a fixed n so query difficulty doesn’t contaminate the comparison; starvation across the seven positions of an n=7 query runs 1.3%, 21.3%, 30.7%, 48.0%, 48.0%, 32.0%, and 13.3%. That is a serial-position curve. The packer protects what you asked for first, protects what you asked for last a little less, and drops the middle. The fragment-scored floor flattens it to 4.0%, 9.3%, 6.7%, 8.0%, 22.7%, 5.3%, and 6.7%.

“A worse retriever with a better allocator wins.”

Topical distance: Sub-intents that span distinct topics starve about twice as often as sub-intents drawn from one topic, at 38.1% against 18.3% for n=7. I predicted the opposite. A topically coherent query gives the reranker a coherent target, and it scores all the correct passages similarly. In contrast, a scattered query lets it latch onto some topics and abandon others.

What the user actually sees

Everything above is retrieval-side. What decides whether any of it matters is what reaches the person who wrote the message, so I generated real support replies from 80 packed contexts and had every reply graded per sub-intent, with both the generation and the grading blind to which arm produced which context.

When the correct evidence reached the packed context, the reply addressed that question 100% of the time, across 261 out of 261 cases, in both arms. Coverage predicts the generated outcome exactly, which is the strongest justification I have for measuring it.

When a sub-intent was starved, the reply answered it anyway 48.1% of the time, based on whatever else happened to be in the window. It explicitly flagged the gap 45.6% of the time, with some version of “I’ll follow up on that separately.” It went silent only 6.3% of the time.

“Starvation mostly does not produce silence; it produces unsupported answers.”

I expected silence, and I was wrong. Starvation mostly does not produce silence; it produces unsupported answers. Whether those answers are actually incorrect is the next experiment, because this harness measures whether a question was addressed, not whether the answer was right.

Some limits: composed queries are cleaner than real support messages, which carry pronouns, implicit context, and conditional clauses. This is one corpus and one embedding family. I drafted the gold labels with model assistance and verified them myself. The model writing those replies was strong, so a cheaper production model would plausibly flag fewer gaps and invent more.

What an allocator actually looks like

Give every sub-intent a floor, and choose it by fragment score. Not the naive floor, which satisfies 99% of its reservations and buys seven points. Select a reservation by relevance to the sub-intent it protects, not by relevance to the message as a whole.

Rerank against the fragment rather than the original query. This inverts what I expected and what I have seen recommended. Scores from different fragments are not comparable across sub-intents, and that incomparability is doing useful work.

Don’t spend your effort on deduplication; near-duplicates cost 1.0% of the budget here. Dedup is worth doing, but it isn’t why your fifth question went unanswered, and treating it as the fix will cost you weeks.

Log per-sub-intent coverage: You already computed it to pack, and it predicts the generated outcome perfectly. A sub-intent that received zero passages is the best predictor available that your reply is about to assert something you cannot support.

Where parallel decomposition breaks

“If it’s late can I get a refund” is one clause and two intents, and the second one’s retrieval target depends on the first one’s answer. Parallel decomposition treats them as siblings. It retrieves the late-delivery policy and the general refund policy, packs both, and misses that the passage you actually need covers refunds for late delivery, which may match neither sub-query particularly well.

There are two ways out: You can tag dependencies at decomposition time, or run a deferred second pass that re-retrieves conditional clauses once the first round resolves.

I would take dependency tagging, for three reasons: A second pass costs a full retrieval round trip inside a latency budget a support bot does not have. The tag is reusable, because a dependent sub-intent should not hold a floor reservation. At the same time, its parent is unsatisfied, so it feeds the allocator directly instead of bolting on a separate mechanism. And it fails visibly, since an untagged dependency shows up as a starved sub-intent in the coverage signal. In contrast, a deferred pass that resolves the wrong condition produces a confident wrong answer with nothing to flag it.

The cost is real; dependency tagging pushes work onto the decomposer, which is already the weakest component in the chain, and I have not measured tagged against untagged. That is a design position rather than a result, and it is the one thing here I am asking you to take on argument instead of evidence.

What to measure on Monday

Take your production pipeline and compute one number: your context budget divided by the average count of distinct questions per incoming message. If that number lands below roughly 1,000 tokens, your allocation policy costs more than your retriever does, and the reranker upgrade sitting in your backlog will buy you less than reserving one slot per question.

On my corpus, the retriever upgrade was worth 12.7 points of coverage, and the allocation change was worth 20.5, which is why I think the ordering is wrong in most pipelines I’ve seen. That ordering is the falsifiable part. Run the same two comparisons against your own corpus, and if the retriever wins, I want to see the numbers, because that result would tell me the crossover sits somewhere other than where I measured it.

The cheaper thing to do first takes an afternoon. Log, for every multi-intent request, how many sub-intents ended up with zero passages in the packed context. A support system that cannot tell you which question it dropped will keep answering that question anyway, about half the time, out of whatever else was in the window.

The post Query decomposition doesn’t fix context starvation — it just moves it appeared first on The New Stack.

Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers

Q.ANT, a startup out of Stuttgart, Germany, builds processors that use light instead of electricity to do some of the math behind AI. The company pitches them as a way to run AI on a fraction of the power today’s chips need.

Now developers can start writing software for those chips without owning one. Q.ANT pushed a free, open-source software kit to GitHub this week that lets developers build and test programs on a normal computer, then run them on the real chips once they get access.

This is a move out of Nvidia’s playbook. Nvidia owes its lead in AI as much to CUDA, the software developers use to program its GPUs, as it does to the chips themselves. 

But with Q.ANT, the catch is the hardware. Q.ANT’s chips are running at a few research computing centers, and everyone else has to wait “the coming months” for cloud access through German provider IONOS or an on-site server from Q.ANT.

The kit, called the Q.ANT Native Computing Toolkit, is free on GitHub under a license that allows commercial use. Developers can work in Python or C. The key piece is a simulator that mimics the chip on a regular computer, with no Q.ANT drivers required.

What can it do today? The AI tools in this first version focus on running models that have already been trained. The examples read handwritten numbers, identify objects in photos and outline shapes in images. Training still happens on regular CPUs and GPUs.

The pitch for photonic computing is power. AI chips burn a lot of energy moving data back and forth between memory and the processor. Q.ANT’s chips do part of the math with light, specifically wave-shaped functions similar to a cosine, which regular chips calculate digitally. Q.ANT says AI models built around those functions get better results with fewer parameters, the settings a model learns during training. Fewer parameters means a smaller model, less data to move and less power. The kit includes examples comparing a standard model with one built Q.ANT’s way. Those comparisons are the company’s own.

“An ecosystem isn’t created by hardware alone. It emerges when the software layer is open and others can build on it,” said Michael Förtsch, Q.ANT’s founder and CEO. He calls the release the “Linux moment” of photonic computing.

Q.ANT is betting light can do the math itself. Lightmatter, one of the best-known companies in the field, now puts its focus on Passage, which uses light to move data between chips. The idea of light-based AI isn’t new, either. TNS covered MIT’s photonic processor for building optical neural networks back in 2017.

Q.ANT raised €62 million in July 2025 in a round led by Cherry Ventures, UVC Partners and imec.xpand. In March, it said its second-generation chips were running at the Leibniz Supercomputing Centre near Munich. The results it published from there compare the new chip with its old one: more than 50 times faster at the kind of math that does most of the work in AI models, and six times less energy on typical jobs, by the company’s numbers. Its bigger claims, like up to 30 times better energy efficiency, don’t say what they’re measured against.

Good software alone won’t carry a new chip. Nvidia has been building CUDA for nearly 20 years and is still adding to it, including deeper native Python support last year. Graphcore, the British AI chip startup, had its own software kit and still ended up being sold to SoftBank in 2024.

Q.ANT calls this the first openly available software kit for programming a photonic processor. That depends on how you count. Xanadu has offered free, open software for its light-based quantum computers since 2018. For now, developers can play with the simulator. What they can’t do yet is test Q.ANT’s power-saving claims on their own models. That has to wait until the chips open up.

The post Q.ANT gives away the software for its light-powered AI chips in a CUDA-style bet on developers appeared first on The New Stack.

A third option is emerging in the fight over AI and your data

Split-screen video interview with The New Stack host Alex Wilhelm and VAST Data cofounder Jeff Denworth.

Not your keys, not your coins. Not your model, not your data?

Over the summer, the tech industry was consumed by a debate about AI use in the enterprise and the need to protect IP. If an enterprise used proprietary models, was data leakage a necessary evil?

Companies seemed to have two options: They could use state-of-the-art, proprietary models and risk losing control of their data, or they could use open-weight models and never kiss the frontier.

Thankfully, a third option is emerging.

Consider the concern: Company A wants to use LLM B from AI Lab C, and they want to avoid training AI Lab C how to eat Company A’s lunch by building its capabilities into LLM B. A good way to resolve the tension would be to let Company A run LLM B on its own infrastructure, so there’s no risk of its information fleeing on the wind.

AI agents are “creating a whole different set of requirements at the data layer.”
–Vast Data co-founder Jeff Denworth

But that raises another problem: AI Lab C doesn’t want to allow Company A to run LLM B on its own GPUs because it doesn’t want to hand over its model weights. It’s the same IP issue the company ran into, in reverse. You have to solve the trust problem in both directions!

Enter VAST Data co-founder Jeff Denworth and a new product called DataEnclave, which aims to let AI labs and enterprise-scale companies deploy proprietary models in secure compute environments without risking data transfer in either direction. (DataEnclave uses Nvidia’s Confidential Computing technology to make the system tick; Vast Data’s core product is AI OS, infrastructure that fits beneath a company’s AI applications.) 

The New Stack had Denworth on the podcast to chat about the confidential computing market. I was curious about timing. Why did Vast build DataEnclave now? Nvidia began rolling out Confidential Computing in a serious way in 2024, after all. Denworth argues that the market needed the core technology, yes, but also demand.

And until late 2025, AI demand was modest compared to today’s token totals. Once agentic coding tools took off, corporate demand for AI products soared. This led to the pricing crisis we saw in early 2026, and the secure AI usage debate we endured over the summer. 

Performance drove demand, demand drove usage, and usage dug up fresh problems to solve. Now the question for the market is whether or not DataEnclave has solved enough concerns on both sides of the proprietary AI-proprietary data equation. The market will sort that out as it moves through early access and into general availability.

Our conversation goes deep into the arc of AI, where companies are in their AI journey today, and how much data remains to be unlocked inside the enterprise. If you want to feel the acceleration, it’s a fun one!

The post A third option is emerging in the fight over AI and your data appeared first on The New Stack.

Enterprise AI desperately needs to protect data and models. Here’s how confidential AI could do it.

Yellow fingerprints arranged in three rows on an orange background, with several prints partially faded.

Most people already understand what generative AI can do. But enterprises run into problems when they need to give a model access to information that cannot leave their own environment, such as a patient record, a customer’s financial details, or a company’s most valuable intellectual property.

Sending that data to a cloud or SaaS service means it crosses external networks and is processed on infrastructure run by another organization, creating additional concerns about control, accountability, and exposure. That’s where AI enthusiasm collides with production realities. Despite its productivity potential, enterprise AI still faces a fundamental gap in trust and control.

Organizations need to know whether a system will expose information it should protect, act as intended, meet security and performance requirements, and behave safely at machine speed.

Alon Horev, CTO and co-founder of AI operating system company VAST Data, tells The New Stack that the challenge is particularly acute when AI systems handle sensitive customer information. “Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”

“Even if you ask the model today to obfuscate a conversation or redact PII from a conversation, it’s hard to have 100% confidence that’s the case, and that it worked.”

Consider a customer support agent that needs access to an individual’s profile to provide a useful, personalized answer. The organization must ensure that information isn’t exposed to another customer, while also considering whether those conversations can be used for training or system improvement. They might contain personally identifiable information (PII) or other protected details, and the consequences of mishandling them ultimately fall on the organization and the people whose information it holds.

Confidential AI architectures: solving a two-sided trust problem

Enterprise AI has two parties to satisfy: organizations must keep sensitive data under their control, while model builders need to protect the weights and software that represent substantial investments in research, engineering, and IP. They’re understandably reluctant to place those assets in environments where customers, infrastructure operators, or attackers might gain access. That mutual need for control has created a stalemate. How can organizations bring advanced models to sensitive data without asking either side to surrender control?

Horev has seen that the most capable models are increasingly delivered as SaaS services, because that’s the simplest way for their creators to distribute and protect them. Even when a provider offers compliance controls, the enterprise might still shoulder the consequences of a breach, misuse, or regulatory violation. Sending information across the WAN also places it in the hands of more systems, connections, and operators, increasing the number of points that must be trusted and governed. Organizations may also be unwilling, or legally unable, to rely on a provider’s assurances that it will not retain, reuse, or expose their data beyond the intended service. 

For organizations in regulated or data sovereignty-sensitive sectors, that could be an unacceptable trade-off. “Naturally, many organizations are adopting a hybrid strategy,” Horev tells The New Stack. “Some applications and datasets can go to the cloud, while others must remain on-premises, sometimes even in the building, or in the country.”

This is where confidential AI comes in. Encryption at rest and in transit protects data while it’s stored or moving between systems. Confidential computing extends that protection into the processing environment, using hardware-isolated execution to create a protected enclave in which the data and model weights can remain encrypted until they’re released to an approved workload.

Cryptographic attestation verifies the hardware, virtual machine (VM), software, and configuration requesting access before releasing keys. The model builder can encrypt its model using the public key of a specific confidential VM. Only that VM’s corresponding private key can decrypt it within protected memory, enabling the customer to use the model without accessing its weights.

Independent key control preserves the separation between the two sides. The enterprise retains control of the keys governing its data, while the model builder retains control of the keys governing its model. While the workload is running, the infrastructure operator doesn’t control either set of keys.

As AI becomes more agentic, those controls will matter more. Agents will need to access more data, systems and tools, and might act on that information with far less human intervention.

“This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do.”

Those that can’t establish strong privacy and governance assurances for today’s models will find it even harder to deploy agents safely in the future. “This world of agentic AI is moving extremely fast, and we need to limit what an agent can see and do,” Horev tells The New Stack.

From architecture to ecosystem

Many businesses simply cannot manage the integration, security, and maintenance of the entire AI stack, because it requires working separately with each model provider to engineer something that suits both parties. Turning confidential AI architecture into something organizations can deploy is the challenge VAST DataEnclave, which was launched on September 22, intends to address.

As a capability of the VAST AI Operating System, the goal is to bring the model, application layer, and data platform together under customer-controlled operating conditions. The architecture is designed to protect both sides of the equation: the enterprise’s data and the model builder’s weights. The customer retains control of its infrastructure and data keys, while the model provider can make its software available without handing over the underlying intellectual property.

“We’re trying to close the trust and control gap by working with world-class model builders such as Cohere, Deepgram, Factory, Fundamental and TwelveLabs, who continue to innovate and build their expertise,” says Horev. The ecosystem also includes infrastructure and security providers such as Nvidia, CrowdStrike, Fortanix, Nscale, Cisco, and Supermicro. The range reflects the practical challenge: confidential AI needs more than a protected GPU. It requires models, applications, accelerated hardware, data infrastructure, and operational support to work together.

That control also changes the cost conversation, without automatically making AI cheaper. Hosted models can make budgets harder to predict as token consumption varies with usage patterns, agent loops, model architecture, and workload volume. Customer-controlled infrastructure gives enterprises a more defined capacity and cost base: they can plan around GPU clusters they own or have already budgeted for, instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.

“…instead of allowing inefficient model choices or uncontrolled agent activity to generate an open-ended token bill.”

The cluster also imposes a natural ceiling on throughput, which helps organizations understand how much work their infrastructure can handle within a given period. Model providers can then price access by token, task, or license, while the enterprise retains greater visibility into its total operating cost.

Why the data platform is paramount

Confidential AI protects data and model weights during inference, but it’s only part of the production challenge. Real-world AI systems are living environments in which data moves between storage, databases, GPUs, networks, applications, and agents.

That’s why confidential AI can’t be bolted onto a fragmented stack. Businesses need to protect the model, the data, and the infrastructure connecting them as one system. As Horev says: “You need to build security in multiple layers of the platform,” with someone accountable for rapidly updating compromised components.

Confidentiality is only useful if the resulting system can also be operated, monitored, and improved. As AI infrastructure becomes more distributed, it becomes harder to tell what’s happening when something goes wrong and where the fault lies.

Horev recommends a “single pane of glass” across storage, networking, and compute, so teams can see what’s happening and keep resolution times low. If a network port is intermittently failing in a data center, for example, an agent could help identify the root cause, provided it has access to the right operational data and tightly controlled permissions. Those permissions should govern the infrastructure it can inspect, the data it can retrieve, and the actions it can take.

The same applies to monitoring AI workloads. Teams need visibility into performance, failures, and access patterns without exposing the customer data or model weights. Agent sandboxes can limit the systems and tools an agent can reach, while data platform observability can log which data it accessed, what it did with that data, and how it interacted with downstream systems.

Evaluation, therefore, becomes part of production discipline. Teams must observe systems, measure behavior, govern access, and manage change in ways that demonstrate progress. Confidentiality, data-level policy, observability and correctness have to work together.

The emerging ecosystem suggests demand for models that can run securely under customer control, wherever sensitive data resides. These are “living systems,” says Horev. “It’s not just leveraging a feature inside of a wider platform.” 

Visit the VAST Data Confidential AI solution page to learn more about the architecture, ecosystem, and availability.

The post Enterprise AI desperately needs to protect data and models. Here’s how confidential AI could do it. appeared first on The New Stack.

Claude Opus 5.5 wants to finish your coding tasks, not just start them

Anthropic wants developers to use Claude and its family of tools to handle complete coding tasks. The company introduced Claude Opus 5.5 on Tuesday to span more of the software application development lifecycle: from design specification creation through debugging to code generation and testing.

As the first release in a new family of Claude 5.5 models, Anthropic says that Claude Opus 5.5 performs at the level of Claude Fable 5.1 on “most work” tasks and costs around 40% less to run than Opus 5. 

GitHub chief product officer Mario Rodriguez is quoted in Anthropic’s launch announcement on exactly where software engineers sit today with frontier models for code automation. He said that developers want agents.

“In our testing [of Claude Opus 5.5] across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable,” says Rodriguez.

Where does Claude Opus 5.5 get its power from?

Anthropic explains that the Claude 5.5 family’s expanded full-lifecycle capabilities were developed under an established set of practices.

These practice elements include extensive alignment testing (model evaluation processes put in place to make sure actions and outputs closely match human values and intended goals), pre-release evaluation by outside organizations, and safeguards for high-risk areas such as cybersecurity and biology.

The organization says that on its most comprehensive alignment test, Opus 5.5 is the strongest-performing model tested to date, with particular improvements in several behaviors that contributed to recent cybersecurity incidents (e.g., biased reasoning, attempting to escape a sandbox, and others).

Independent SRE & AI reliability architect Akash Thakur tells The New Stack that Anthropic’s work getting its model to complete whole coding tasks is impressive, but “getting it to know when it hasn’t” is the harder problem developers need to think about.

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces — and that’s the real unlock, because momentum on software comes from finishing things, not starting them,” Thakur says. “…But ‘completed’ and ‘correct’ aren’t the same thing. The task that looks done and completed is exactly the one that often costs the team later, so the win isn’t removing the human — it’s moving them from writing the code to verifying it.”

“Models working at the level of Claude Opus 5.5 are genuinely good at breaking a project into small, finishable pieces …but ‘completed’ and ‘correct’ aren’t the same thing.”

Suggesting that we are now witnessing a “higher bar for more capable models”, Anthropic said that models that could fully automate AI research itself should meet a higher safety standard. The company noted that as AI becomes more capable, public policy should play a larger role in making sure these systems are safe. Its recent work with Accenture is offered as an example of how Anthropic is building the infrastructure to support this.

Sprawling jobs: codebase-wide migrations & audits

Getting more specific, Anthropic has claimed Opus 5.5 is “particularly good” at long, sprawling jobs like codebase-wide migrations and audits. An early tester said it audited and fixed a 200,000-line codebase in under three hours, while Opus 5 took over 20 hours and used 2.5x as many tokens. 

“In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less,” said Anthropic in a press statement.

Field CTO for the EMEA region at Coder, Eric Paulsen, tells The New Stack that he is happy to hear Claude Opus 5.5 is becoming more holistically capable, but he’s “not surprised” because AI is “eating the software delivery chain end to end” today.

“Despite the efficiencies showcased in Opus 5.5, the moment an agent can reliably finish real work, the question stops being whether the model is good enough and becomes where you’re letting it run,” Paulsen says. 

“Kicking off a Claude Code session on a laptop that can be compromised, stolen, or simply run out of compute is not an engineering environment; capable agents need dedicated, governed infrastructure with the same guardrails, secrets handling, and observability a developer would demand of any other production workload. As we’ve seen here with Anthropic’s own work, firms need to underline testing and safeguard procedures for launches of this kind,” he added. 

Anthropic assures users that external testing has been conducted and that the model was tested before release by Frontier Design and METR. When established safeguards for Opus 5.5 intervene, requests “fall back to another model transparently,” meaning developers may not see which model actually handled a given call.

In cybersecurity, users can identify and fix bugs in their code, but most cybersecurity tasks will be re-routed to Opus 4.8. Requests flagged by biology and frontier LLM development classifiers will be re-routed to Opus 5. Vetted organizations can apply to Anthropic’s Life Sciences Verification Program to use Opus 5.5 for biology research, and the company says it will expand its Cyber Verification Program in the coming weeks.

The developer’s terminal prompt has fundamentally changed

HasData co-founder, Sergey Ermakovich, tells The New Stack that his work as a web scraping specialist for data pipelines and AI means he sees zen-like, one-brick-at-a-time logic in what Anthropic has done. 

“The biggest change — and it’s a trend that will have driven Anthropic’s design and development aspirations for Claude Opus 5.5 — is that the terminal is no longer just a place where a developer pastes generated code,” Ermakovich says. “The terminal now becomes part of the model’s workspace.”

Because a model can run commands, inspect failures, modify files, and verify the result, Ermakovich suggests it can “close the loop” instead of handing unfinished work back to an engineer. “That makes small complete tasks much more valuable. Fix one failing test, update one dependency, migrate one endpoint, verify it, then move to the next task,” he adds.

“The terminal is no longer just a place where a developer pastes generated code — the terminal now becomes part of the model’s workspace.”

Commenting on the development of this model as part of Anthropic’s approved corporate messaging, John Ruelas, staff software engineer at Ramp, said that verbose, hard-to-follow output has been his “biggest frustration” with frontier models, but “Claude Opus 5.5 fixes it” for him.

He noted that it writes like a good colleague and follows his company’s writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of his prompts, he preferred its version to his own. When it optimized the team’s test suite, he could follow its reasoning easily and “shipped the change with confidence.”

We’re so done with the autocomplete era

Vice president of AI strategy at Abbyy, Maxime Vermeir, tells The New Stack that the more lifecycle-wide Claude Opus 5.5 features on offer show that “we’re done with the autocomplete era” for basic code automation tools.

“Anthropic’s elevation here reflects the fact that it used to be thought of as marvelous if AI could complete a developer’s next line of code, but today the expectation is that you hand it a whole Jira ticket and that it gets done,” Vermeir says. “But despite the power on show with Claude Opus 5.5, the question remains as to how much these newer models will actually understand what ‘done’ means, as often it seems they have a desire to keep burning tokens by offering you yet another thing it didn’t quite do right.”

Opus 5.5 also “communicates more naturally” than prior models, with early testers finding its writing clearer and easier to follow, making it a better work partner over long sessions.

Costed out lower than Opus 5, Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens (20% less than Opus 5), and Anthropic has cut cache read prices by 60% for token-billed usage. Opus 5.5 also needs fewer tokens for higher-quality work and generates output more than 30% faster than Opus 5.

The post Claude Opus 5.5 wants to finish your coding tasks, not just start them appeared first on The New Stack.

Your AI agent is burning tokens on choices that don’t need words

Blur or abstract motion

AI agents spend a ridiculous amount of compute generating text nobody actually needs. The decisions an agent makes along the way don’t require a written answer and yet, agents still send them to generative models, wait for an answer while burning through tokens and then parse that output back. The overhead is already drawing scrutiny — OpenAI’s own researchers recently disclosed spending $7,000 a day running agent workloads.

Kev, a new family of open decision models built on Qwen 3.5, takes a different approach and skips the generation entirely.

Developer Jared Palmer released a new generation of Kev on Sunday, with 0.8 billion, 4 billion, and 9 billion parameter models built on Qwen 3.5. Kev is prefill-only, processing the state, questions, and candidates in a single forward pass before reading the decisions from a pointer head without an autoregressive decoding loop.

Kev, a new family of open decision models built on Qwen 3.5, takes a different approach and skips the generation entirely.

Decisions without generated text

Kev supports three decision types: Noul for yes/no, Choice for selecting among candidates, and Score for ordered levels, mirroring TypeSafe’s System One API. Developers provide the state and questions, and the pointer head returns probabilities across the available candidates.

For a tool-routing decision, the output could look like this:

search: 0.82

database: 0.13

calculator: 0.05

Kev can still choose the wrong tool, but because it scores only the candidates it’s given, it can’t introduce an option that isn’t on the list.

Routing, safety checks, escalation, and ranking can then move to the decision layer, leaving larger reasoning models to handle the open-ended work.

Kev can still choose the wrong tool, but because it scores only the candidates it’s given, it can’t introduce an option that isn’t on the list.

Batching choices, one pass

Multiple decisions can also be made against the same state in a single forward pass, with a block-causal attention mask isolating the questions while the pointer head scores each set of candidates independently.

Palmer’s documentation shows the 4B model processing three questions in 277 milliseconds in bf16 on an M5, although without a controlled comparison against Qwen generating equivalent answers on the same hardware, the result doesn’t establish how much faster the approach is in practice.

The ability to evaluate several decisions against the same context could become more useful as agent loops grow more complex, but skipping generation doesn’t make the resulting decisions inherently better.

Calibration limits and tradeoffs

The largest model, Kev-9B, reached 83.7% accuracy on the project’s locked out-of-domain test, according to Palmer’s model card. That’s a developer-reported benchmark, and Palmer documents some limitations alongside it.

The probabilities Kev returns don’t always reflect how confident developers should be in the result. Palmer found that temperature calibration can drift on unseen source distributions, a problem for agents that use probability thresholds to decide whether to execute an action or escalate it, since even a high-probability choice can still be wrong.

Fine-tuning also changes some of the capabilities inherited from the underlying model. Palmer’s evaluations show declines on general-knowledge and arithmetic tests, particularly among the smaller models. That’s consistent with Kev’s more specialized role alongside a general-purpose model, although its performance in dynamic agent environments will also depend on how well it handles tools, choices, and labels it never encountered during training — and debugging agent failures often points to infrastructure rather than the model itself.

The approach predates Kev. TypeSafe introduced Jev earlier this month as part of its System One platform, using the same Noul, Choice and Score primitives, and Kev implements its /v1/systemone request and response format so applications built against the API can point to a local Kev server instead.

Open weights, open training

The biggest difference is that Palmer released Kev under Apache 2.0 with the model weights, training code, and evaluation tooling, giving developers the option to run and train it on their own infrastructure. Jev’s weights and training data aren’t public, however, which makes direct performance comparisons difficult because differences between the models can’t be isolated to architecture, size, or training.

For applications that make only a handful of bounded decisions, constrained decoding on a model that’s already running may be simpler than adding another model to the stack. Agent loops can make those decisions constantly, however, moving through routing, ranking, safety checks, tool selection, and escalation before generating much user-facing text. It’s a pattern showing up across model architectures — stripping out unnecessary computation when the task doesn’t require it.

When those steps only require a choice or probability, Kev can handle the decision directly while leaving open-ended reasoning and final responses to the larger generative model.

When those steps only require a choice or probability, Kev can handle the decision directly while leaving open-ended reasoning and final responses to the larger generative model.

The post Your AI agent is burning tokens on choices that don’t need words appeared first on The New Stack.

Kubernetes can run AI inference. But can it count the real cost?

Abstract 3D illustration of interconnected purple geometric nodes and gold lines representing a distributed network or cloud infrastructure.

Welcome to another edition of Road to KubeCon, where we’re tracking the Kubernetes and cloud-native ecosystem on the way into KubeCon + Cloud Native Con NA 2026, to be held in Salt Lake City, Utah, November 9-12.

This week, we look back at the past week of significant movements in the Kubernetes space. Most notably, we see interesting advances in cloud-native architectures for AI inference. We take a look at that, plus a new Gartner quadrant, new Kubernetes hardening updates, and important CNCF project updates.

HPE challenges server virtualization platforms

On Monday, Gartner published its Magic Quadrant for Server Virtualization Platforms, a guide comparing solution providers in the server virtualization market. The quadrant names HPE as a Challenger based on Ability to Execute and Completeness of Vision.

Hewlett Packard Enterprise (HPE) is a presenting sponsor of Road to KubeCon. HPE Software helps IT organizations modernize infrastructure, streamline operations, and accelerate AI initiatives across hybrid, multi-vendor environments.

According to the HPE newsroom, the recognition reflects ongoing momentum behind HPE Morpheus Software, its virtualization and cloud operations portfolio. HPE was positioned in the Challengers quadrant alongside Canonical and Oracle.

The announcement comes as enterprises rethink their virtualization strategies. Rather than simply swapping in another hypervisor, HPE argues that organizations increasingly need unified governance and ways to provision, orchestrate, observe and secure workloads — including VMs, containers and AI workloads — across clouds.

Kubernetes hardens container storage

On Wednesday, Red Hat’s Nispriha Jagan and Neeraj Krishna wrote on the Kubernetes project blog about two new storage security features shipped as Alpha in Kubernetes v1.37, which included 67 enhancements.

The notable security features are new bind mount options and emptyDir permissions. The enhancements come as multiple security findings have surfaced regarding emptyDir volumes, one of the most common writable volume types. The additions are made possible by low-level Linux security mechanisms.

According to the authors, these enhancements give users native controls to harden Kubernetes workload storage better. “Supporting noexec, nodev, and nosuid gives users a native way to harden volume mounts to match security benchmarks and policy,” the authors write.

KubeCon adds AI Inference + Agentic track

Last month, CNCF announced it will feature an AI Inference + Agentic track at KubeCon + CloudNativeCon North America 2026, exploring the intersection of generative AI and cloud native infrastructure. Attendees can explore the sessions here.

The added track underscores the growing use of Kubernetes for production AI workloads, particularly as the focus shifts from training models to serving them in production. It also reflects emerging practices for building agentic systems around protocols like MCP and A2A, as well as infrastructure such as AI gateways.

China Merchants Bank unifies AI inference on Kubernetes

China Merchants Bank, a leading Chinese commercial bank, recently showcased its cloud-native AI infrastructure at a CNCF event in China. Its infrastructure team won the CNCF End User Case Study Contest with an architecture combining Kubernetes with several cloud native projects:

  • Kueue, for job queueing and quotas,
  • KEDA, for event-based auto-scaling,
  • Prometheus, for systems monitoring and metrics,
  • HAMi, for sharing accelerator capacity across Kubernetes workloads,
  • and Fluid, for accelerating access to datasets.

The bank has a large pool of nearly 10,000 accelerator cards used for AI computation. These are heterogeneous, meaning they are not all the same type or configuration.

According to the CNCF announcement, the architecture unified management of 99% of its AI compute resources, while increasing average utilization from 35% to more than 60%. It also cut the cost of processing 1 million tokens by 60% under comparable conditions.

The case study shows how cloud-native infrastructure can improve utilization and efficiency for AI training and inference, even in regulated areas like financial services.

Industry take: Can AI inference on Kubernetes handle token cost issues?

Interest in AI inference on cloud native infrastructure is palpable. However, this week Val Bercovici, chief AI officer at WEKA, an AI-native data platform, questions whether Kubernetes’ existing resource model fits the changing economics of large-scale AI inference.

Bercovici tells The New Stack: “With AI inference, it’s cost per token, and that cost depends on state Kubernetes was never designed to manage: request mix, KV cache occupancy, the balance of prefill and decode, and how memory and bandwidth are consumed inside the accelerator after a pod is already running.”

“My view is that Kubernetes doesn’t go away,” Bercovici says. “But unless its resource model evolves, it becomes a tax on inference economics.” 

He foresees a new scheduling and memory layer to emerge around Kubernetes that can compute what a token actually costs to serve. Then platforms could make more informed, cost-based decisions about how inference workloads are scheduled and served.

As Kubernetes evolves, so do the demands on the teams running it. Presenting sponsor HPE helps teams address that complexity with software spanning virtualization, cloud management, observability, and automation.

Move over, platform engineering. Hey, agentic engineering.

A new Weave Intelligence report, State of AI in Platform Engineering Volume 2, authored by Sam Barlien, Luca Galante, and Florian Lipp, surveyed 242 platform engineering leaders on the before-and-after effects of introducing agentic AI into platform engineering.

38% of teams are shipping at least twice as much as before AI. When assessing ROI across the software delivery life cycle, 20% report efficiency gains and 11% report operational savings. Yet only 8% report a transformative, structural shift. Meanwhile, 29% are still prototyping without realized gains, with some outliers reporting negative results.

The biggest roadblock to scaling AI usage? A lack of platform readiness, including APIs, deterministic pathways, and standardization. Weave’s takeaway is that platform engineering must increasingly account for AI readiness and agentic experience as agents become another key platform consumer.

OpenTelemetry Kubernetes attributes processor reaches v1.0.0

On Wednesday, OpenTelemetry, the graduated CNCF project and open standard for telemetry, announced the v1.0.0 release and distribution of its Kubernetes attributes processor. It’s a helpful feature that uses the Kubernetes API to add Kubernetes metadata, such as stability, distributions, warnings, issues, and other metrics, to resource attributes.

According to the release notes, written by Elastic’s Christos Markou and Datadog’s Pablo Baeyens, the feature has been in progress in the OpenTelemetry Collector SIG since late 2025, based on a roadmap of users’ most-requested features. Existing attribute processors should review the breaking changes and migration guide.

DigitalOcean opens Spot GPU node pools

Technically, this occurred the week before last, but didn’t make the digest. As of September 9, DigitalOcean Kubernetes’ (DOKS) Spot GPU Node Pools entered public preview. According to the release notes, the feature runs worker nodes on interruptible GPU capacity at a “lower, variable rate than on-demand GPU nodes.” This could offer a cost-effective option for fault-tolerant workloads.

Other updates from the K8s universe

More updates from the infrastructure-heads, platform engineers, and multi-cloud operators working in the Kubernetes ecosystem: 

Follow the Road to KubeCon

Road to KubeCon is an eight-part series presented by HPE, which will be at KubeCon + CloudNativeCon North America in Salt Lake City. Before you go, explore how HPE Software helps IT teams do more with less complexity.

We’ll be here every Friday until KubeCon.

If you’d like to participate, Bill Doerrfeld, the writer of this series, is open to pitches — you can send release notes, quotes, reports, videos, case studies, or hot takes through his contact page.

If you didn’t catch the inaugural edition covering the Kubernetes v1.37 release, check it out here. You can also visit the Road to KubeCon page for the complete archive.

The post Kubernetes can run AI inference. But can it count the real cost? appeared first on The New Stack.

Your agent is only as good as your infrastructure

Dark abstract 3D glass ribbon rendering symbolizing complex AI agent infrastructure and bursty data workflows.

You built a great agent, but something happened when it moved into production.

In testing, your agent reviewed pull requests efficiently on its own. It read the diff, grepped the codebase for related usages, ran the test suite, checked whether CI was still red from an earlier commit, and drafted a comment—all before you’d finished reading the diff yourself.

In production, however, imagine the same five steps ran behind every other PR review your team’s agents performed that hour. Some reviews landed in seconds; others sat for minutes because the test-suite step landed on a node that was mid-burst from someone else’s agent. 

The agent didn’t change. The execution environment did, and that’s what decided whether review time held steady or crept up.

Agentic applications introduce a different execution pattern than traditional chat applications. As those workflows become longer and more dynamic, infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.

That’s why your agent is only as good as your infrastructure.

Agents aren’t chatbots with more steps

The difference between serving inference for a chatbot vs. an AI agent isn’t simply that one is “more capable.” They execute work differently:

  • A chatbot usually makes one inference call per user message. The model receives a prompt, generates a response, and waits for the next user input before proceeding.
  • An agent executes the entire workflow, turning one user message into a chain of inference calls. It might decide to search documentation, retrieve data from a database, call an API, execute code, evaluate the result, and then repeat that process before producing an answer. Each of those decisions may trigger another inference call, and every result becomes additional context for the next step.

“Infrastructure has a much larger influence on latency, reliability, and cost than it does for a simple chatbot.”

That execution model changes the infrastructure requirements for AI agents. 

One question, many steps behind it

Instead of optimizing for individual inference requests, the system has to support long-running workflows whose latency and reliability depend on every component in the chain.

A single user request often expands into a sequence of inference and tool execution steps, sometimes called multi-turn tool calls or the agentic loop. Rather than generating one response, the model alternates between reasoning and interacting with external systems.

For example, you ask an agent why checkout latency spiked overnight. The agent pulls the deploy log, queries the monitoring system, runs a diagnostic against the connection pool, weighs whether the culprit is a bad deploy or a capacity issue, and then folds that into another inference call before finally producing a full-fledged response.

Each reasoning step becomes another inference request, and every tool result is added to the model’s context before the next step.

This workflow changes what reliability means

Multi-turn workflows are inherently sequential, which is why even low latency can quickly add up to a significant amount. Every inference step waits for the previous one to finish. If a database query takes two seconds, the model can’t begin the next reasoning step until that result returns. The model may generate tokens quickly, but the other steps slow it down.

“In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link.”

In this agentic workflow, every step in the chain has to hold, because the chain is only as strong as its slowest link. Instead of processing isolated inference requests, the inference stack has to orchestrate a chain of dependent model invocations and external tool calls. As those workflows become longer, the stack increasingly determines how quickly, reliably, and cost-effectively the application performs.

But the user doesn’t see an orchestration hiccup. They see an agent that hung or gave up.

That’s why end-to-end agent reliability depends on much more than model quality. Infrastructure determines whether each step has the resources it needs to execute predictably under load.

Why the bill and the performance both feel unpredictable

A second difference in agentic workflows catches teams off guard: demand patterns and their impact on your inference bill. 

Most inference services, and the pricing built on top of them, assume traffic arrives at a predictable pace. A typical inference solution knows the predictable demand pattern: User traffic increases, request volume increases, and capacity scales accordingly. Cloud infrastructure is typically optimized for these steady request patterns using mechanisms such as autoscaling, load balancing, and capacity planning.

Agent workloads don’t behave that way. Individual workflows pause while waiting on external systems, then resume as soon as new information becomes available.

  • The pause: The agent waits on an external API or database, so the GPU serving that workflow has no inference work to perform, and its accumulated context may be evicted from GPU memory while it waits
  • The burst: As soon as external systems return results, inference resumes simultaneously across many workflows, creating short and sharp spikes in GPU demand, each re-processing its full accumulated context

If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat — flat, then a spike every time tool results come back — that’s the signature. It means you’re provisioning for the average when you should be provisioning for the peak, and it’s usually the first place p99 latency quietly blows out.

“If you’re watching GPU utilization and it looks less like steady load and more like a heartbeat, that’s the signature.”

Inference services designed around steady or predictable request streams can struggle to allocate resources efficiently under these conditions. This leads to inconsistent latency, GPU underutilization, or higher operating costs.

When performance becomes unpredictable, or a bill doesn’t match what you expected, it’s evidence your infrastructure was built for a different workload than the one you’re actually running.

What infrastructure built for agents actually looks like

Agentic applications place different demands on infrastructure than traditional AI workloads: long dependency chains, bursty demand, and continuous evolution. Because of this, the infrastructure is deciding whether the chain holds, and whether the bill holds too.

For an agent to run, it needs infrastructure purpose-built to support:

Performance that holds across the whole chain. The infrastructure must keep latency consistent across multi-step and multi-tool workflows.

Scalability that responds to bursty demand. Infrastructure should scale quickly as inference demand fluctuates, without requiring capacity to remain provisioned during idle periods.

  • Predictable economics even for dynamic workloads. The infrastructure bill should reflect actual usage.
  • None of this makes bursty demand disappear, but it changes how the system absorbs it. A large enough simultaneous burst, or a workflow that accumulates enough context before pausing, still costs something. The goal isn’t zero cost or zero limit; it’s making both predictable.

Your agent is only as good as your infrastructure. Get it right, and your agent’s responsiveness, reliability, and cost-effectiveness will improve your work. 

Learn more about how infrastructure can be purpose-built for agentic workflows: Check out the documentation to get started.

The post Your agent is only as good as your infrastructure appeared first on The New Stack.

❌