❌

Normal view

The Control Gap: Enterprise AI organizations have an ownership problem, not a technology problem — and most are governing it by hand

1 July 2026 at 21:50

AI portfolios are expanding far faster than the ability to govern them across enterprises. Most organizations run a contested field of platforms, each claiming to be the “primary” AI layer; few could confidently detect a model drifting or failing in production; and the single most-cited barrier to control is the absence of any one owner accountable for AI across the stack. The result is a widening control gap — ambition and spend racing ahead of visibility, ownership, and cost control — with autonomous agents already producing real financial and operational failures.

This wave of VentureBeat Pulse Research examines the enterprise AI control gap: how many platforms claim to be the primary AI layer, who actually governs AI behavior across them, whether organizations could detect a model failing in production, what most blocks cross-platform governance, and how the financial and operational control failures of autonomous agents are already surfacing.

The central finding is a control gap — the distance between how aggressively enterprises are expanding AI and how little of it they can see, own, or govern. Just under three-fifths (58%) are net-adding AI initiatives, with “expanding significantly” the largest single posture.

Yet 85% run two or more platforms each claiming to be the “primary” AI layer and only 8% have consolidated to one. Against that contested surface, 40% say they are very confident they would detect a model drifting, behaving unsafely, or failing in production — but only 10% back that confidence with active monitoring and alerting, the rest leaning on manual human review. The machinery to expand AI is running well ahead of the machinery to control it.

The gap is, above all, a question of ownership. Only a third (38%) say a central team governs AI today, and a fifth (20%) say each platform team governs its own independently; the single most-cited barrier to cross-platform governance is the absence of a single accountable owner (32%), and roughly one in six (17%) say no role holds formal accountability at all. The same vacuum shows up in spend: just under half (49%) name shadow AI — unauthorized agentic pipelines run on corporate cards outside central oversight — as their most severe control failure, and another 25% have been hit by a runaway “infinite loop” agent bill. Enterprises have standardized the ambition well before they have standardized the control.

Methodology

VentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on the enterprise AI control gap — governance, observability, and cost control across multiple AI platforms. Responses are filtered to organizations with 100 or more employees and, for this cut, exclude the respondents who selected “Other” as their job function, leaving a base of identifiable roles (n=145); all are drawn from a single Q2 2026 (June) wave. 

By organization size the sample tilts toward the mid-market and lower-large bands: 100–499 and 500–2,499 employees (23% each) lead, with 10,000–49,999 (22%) and 2,500–9,999 (20%) close behind and 50,000+ at 11%. By role it is senior and technical: consultants and advisors (20%), CIO/CTO/CISO (18%), directors of engineering/IT (14%), product and program managers (13%), and enterprise architects (12%) make up the core. Technology/Software is the largest industry at 41%, followed by Financial Services and Professional Services (12% each) and Healthcare/Life Sciences and Manufacturing/Industrial (10% each).

The findings should be read as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. Where a single share would be fragile on its own, the report leans on the direction and grouping of responses rather than the exact percentage point.

Finding 1: Expansion is outrunning control

AI portfolios are growing faster than the means to govern them

We asked enterprises to describe how their AI portfolio has changed over the past 12 months. Growth leads — with a meaningful minority deliberately pulling back.

Expansion leads. Combining “expanding significantly” (33%) and “net positive growth” (25%), just under three-fifths of enterprises (58%) are net-adding AI initiatives. Yet a substantial share is easing off deliberately: roughly a quarter (23%) are actively rationalizing — scaling what works and cutting the rest — and another 12% hold their portfolios flat. Only a handful (3%) have paused to get governance in order first.

This is the engine behind every gap that follows: enterprises are accelerating into a landscape they have not yet learned to see or own, and a notable 4% cannot even describe their own portfolio. The ambition documented here is exactly what makes the visibility and ownership shortfalls in Findings 3 and 4 consequential rather than academic.

Finding 2: No single “primary” AI layer — the surface is contested

More than four in five run multiple platforms each claiming primacy

We asked how many enterprise platforms currently claim to be the organization’s “primary” AI layer — the ERP, EHR, ITSM, productivity suite, or data platform each positioning itself as the center of gravity. Almost no one has a single answer.

The defining condition is contested primacy. Adding the two multi-platform bands, 85% of enterprises have at least two platforms each asserting itself as the primary AI layer, and more than a third (36%) describe an open four-way-or-more contest. Only 8% have consolidated to a single layer, and another 6% have not even mapped the question. This is the structural reason governance is hard: there is no agreed center of gravity to govern from. Each platform brings its own AI, its own controls, and its own assumptions — and, as Finding 3 shows, the question of who governs across them increasingly has no settled answer.

Finding 3: Governance is claimed at the center but contested in practice

A central team owns it on paper; in practice, it's fragmenting

We asked who is actually responsible for governing AI behavior across all of those platforms today, and which function holds primary accountability. The headline answer is reassuring; the detail is not.

On the surface, a central governance function is the leading answer — but only a third (38%) claim one, well short of a majority. The rest of the distribution undercuts it further: a fifth (21%) say ownership is unclear or contested between teams, a fifth (20%) say each platform team simply governs its own AI independently, and 19% say no one has addressed it at all.

Accountability fragments further when we asked which role actually holds it — CIO/CTO/CISO leads at 27%, a Chief AI Officer or equivalent at 22%, and a striking 17% say no one holds formal accountability yet. Even where a central team is claimed, the named owner is most often the general technology executive rather than a dedicated AI authority. The governance function exists more often as an org-chart aspiration than an operating reality — the precondition for the detection gap in Finding 4.

Finding 4: The detection gap — confidence is real but largely manual

Only one in 10 have active monitoring and alerting

We asked how confident enterprises are that they would detect an AI model in production that was drifting, behaving unsafely, or failing to complete tasks correctly. This is the heart of the control gap.

This is the report’s central number. While 40% say they are very confident they would detect a failing model, the overwhelming majority of that confidence rests on manual human review (30%) rather than automation — just 10% have active monitoring and alerting actually in place.

At the other end, more than a quarter combine the two reactive answers — no systematic visibility (8%) and would hear it from end users first (19%) — meaning they would learn of a production failure after the fact, from the people it affected. The plurality (32%) sit in a hopeful middle, expecting to “catch most issues eventually.” Set against the aggressive expansion of Finding 1, this is the crux of the control gap — enterprises are scaling AI into production faster than they are building automated means to know when it breaks. Confidence is real, but it is largely manual, and automated detection remains the exception.

Finding 5: The missing owner is the biggest barrier

Governance stalls on accountability first, visibility second

We asked enterprises to name their single biggest barrier to governing AI across multiple platforms. The org chart tops the list.

The single missing owner leads at 32%, the most-cited barrier. Vendor opacity (25%) and the lack of tooling or infrastructure to observe across platforms (16%) sit behind, and together these two technical-visibility barriers (41%) outweigh the ownership gap. Leadership deprioritization accounts for another 17%, while a clear lack of talent is rare (5%). Rounding out the picture, another 5% say it isn't a barrier for them at all — they've already solved it.

Read together, the picture is more contested than the headline suggests: enterprises still most often name a missing owner, but a good share locate the obstacle in vendor black boxes and the absence of cross-platform observability.

Asked in a free-text question what one thing they would fix, respondents converged from different directions on the same answer — a single accountable owner, and a control plane that abstracts cost, drift, and model choice away from the end user.

Finding 6: The fine-tuning ROI reckoning

Roughly seven in 10 have little to show for custom model investment

We asked what share of the proprietary foundation models enterprises have invested in fine-tuning over the past 18 months have delivered clear, measurable positive ROI in production today. Most describe a sandbox graveyard — or a deliberate decision to avoid one.

Custom fine-tuning has, for most, not paid off. Combining the three disappointing outcomes — sandbox graveyard, strategic avoidance, and total write-off — roughly seven in ten (73%) either failed to get custom models into productive use or deliberately declined to try, against 27% for whom fine-tuned models are a reliable advantage. The largest single group (45%) remains the graveyard: projects too expensive or complex to maintain, stranded in development. Another quarter (24%) never started — they priced in the downstream maintenance burden and avoided it.

The signal is that many enterprises still treat bespoke model training as a cost trap, which helps explain the pragmatic, buy-and-blend vendor posture in Finding 7.

Finding 7: Vendor posture — hybrid by default, with defection rising

Enterprises blend open and closed models; more are now trimming a vendor

We asked two related questions: whether enterprises are shifting workloads toward open-weight models to escape API costs and lock-in, and which proprietary vendor, if any, they are most likely to phase out over the next year. The answers describe hedging — and a rising willingness to cut.

On open weights, a clear majority (51%) strike a hybrid balance, with a deliberate closed commitment second at 32% and a hard pivot to self-hosted open models at 16%. The hybrid plurality is the same instinct visible throughout this survey — keep optionality, avoid being trapped — while the closed group remains candid that the operational overhead of self-hosting still outweighs the savings for them.

On vendor defection, loyalty by inertia no longer leads: Microsoft is now the single most-named target (29%, often citing Copilot/Azure cutbacks in favor of direct model access), narrowly ahead of the 27% who are downsizing no one at all. OpenAI follows at 21% (citing pricing volatility), with Anthropic at 15% and Google at 6%. No single vendor faces a wholesale exodus, but among identifiable roles the balance has tipped from “expanding across all” toward actively trimming at least one provider.

Finding 8: The agentic spending crisis — shadow AI leads the failures

Unauthorized pipelines, not runaway loops, are the top control failure

Finally, we asked what the most severe financial or operational control failure enterprises have experienced as autonomous agents run over longer execution windows. Shadow AI tops the list — and very few have escaped a scare.

The control gap has a price, and it is being paid. Just under half of enterprises (49%) cite shadow AI — unauthorized agentic pipelines spun up on corporate cards outside any central oversight — as their most severe failure, the operational twin of the “no single owner” barrier in Finding 5. Another 25% have been burned by a runaway infinite-loop agent bill, and 6% by an agent that degraded production databases. Only 21% report guarded stability — the minority that has imposed hard token throttling and budget caps at the infrastructure layer and avoided surprises.

Put differently, roughly four in five of these enterprises (79%) have already experienced a real financial or operational control failure from autonomous AI, not merely worried about one. As with detection in Finding 4, the deterministic controls that would prevent these failures exist at only a fraction of organizations.

The bottom line: A control gap that spending cannot close on its own

Organizations with 100 or more employees describe AI programs that are expanding fast and governing slowly. Just under three-fifths are net-adding to their portfolios; more than four in five run a contested field of platforms with no agreed primary layer; and the thing they most often name as their chief obstacle is a single accountable owner. The visibility to match the ambition is largely manual — only 10% have active monitoring and alerting, and confidence in detecting a failing model rests mostly on human review rather than automation.

The consequences are already concrete rather than hypothetical. Custom fine-tuning has disappointed more often than not, pushing enterprises toward a hedged, hybrid, buy-and-blend model posture; and the autonomous agents now reaching production have produced real control failures for roughly four in five respondents, led by shadow AI running outside any central oversight. This reads as a directional signal rather than a precise measurement — but the direction is consistent across every question: ambition, spend, and deployment are racing ahead of ownership, observability, and cost control. The control gap is not a tooling problem that more spending will close on its own; it is, first, a question of who owns the answer. 


Based on survey responses from 145 qualified enterprise respondents (100+ employees). Sample size is small; data should be treated as directional. Respondents include Directors, VPs, CIOs, CTOs, and Enterprise Architects across Technology, Financial Services, Retail, Healthcare, and other sectors.

T-Mobile moving tens of thousands of virtual machines off VMware amid lawsuit

1 July 2026 at 21:21

T-Mobile is asking a New York court to rule that Broadcom was contractually obligated to continue supporting its VMware perpetual licenses.

In its complaint, T-Mobile said it has tens of thousands of virtual machines using VMware software across approximately 303,140 CPU cores. It also said that it was migrating off VMware but noted the time-consuming and technical challenges involved in migrating over 1,000 applications.

It filed its lawsuit, which was first reported by The Register today, in the Supreme Court of the State of New York in August 2025 (PDF).

Read full article

Comments

© Getty Images | Anna Moneymaker

OpenClaw’s new app doesn’t run AI on your phone. That’s the whole point.

Person holding a smartphone horizontally while photographing white clouds against a blue sky.

OpenClaw finally dropped its iOS and Android apps this week, meaning you can now ditch the Telegram and WhatsApp methods to talk directly to your personal AI agent. But what’s arguably more exciting is that the app isn’t actually running the AI on your phone. It’s just hooking up to an agent you’ve already got running somewhere else. Your phone now acts as a window into that agent, complete with voice, notifications, and camera access.

It’s a nice design choice, and exactly where personal AI agents are headed.

Phones become authenticated endpoints

The phone is basically becoming a really smart remote control for OpenClaw. Instead of cramming an increasingly powerful agent onto a phone with battery and memory constraints, developers are treating the phone as one more screen for an agent that lives elsewhere. The agent keeps working whether your phone is in your hand or charging in the other room.

Within this model, the phone approves actions, pings you with notifications, lets you talk to the agent, and shares your camera when the agent needs eyes on something.

Persistent runtimes replace mobile constraints

But OpenClaw isn’t the first to do this. Anthropic’s Claude Cowork with Dispatch follows a remarkably similar pattern. Users assign work from their phones, but execution occurs on a persistent desktop runtime. The mobile app acts as a companion for starting tasks, monitoring progress, and receiving results rather than becoming the agent itself.

OpenAI is moving in a similar direction as well. With Codex, developers increasingly interact with long-running coding agents that continue working independently and can be checked on from multiple clients, instead of treating the phone as the place where the agent runs.

Different companies, different products, but a similar architectural bet to keep the agent running in a persistent runtime and give people lightweight clients to interact with it.

When multiple teams independently converge on the same architectural pattern, it’s often an early signal that the industry has found a model that solves a real engineering problem.

The engineering problems are totally different now

This shift changes what developers spend their time thinking about. Building mobile apps used to mean worrying about battery life, memory limits, offline mode, and squeezing the best performance out of a phone. If the agent is running somewhere else, most of those concerns fade into the background.

Now, a new set of questions comes to mind, such as how a phone securely connects to a long-running agent? How do you manage permissions across multiple devices? What happens if every client disconnects but the agent keeps working?

Agent identity beyond login screens

There’s a downstream effect here. Once the phone is just one of several trusted endpoints talking to your agent, you need a much more robust approach to identity. You’re not logging a user into an app anymore. You’re authenticating devices into an ongoing relationship with a persistent agent.

As that agent gains the ability to read your files, send emails, call APIs, and control external tools, authentication becomes load-bearing infrastructure.

Distributed agents reshape developer tooling

Zooming out a bit and looking at the bigger picture highlights how personal AI agents increasingly resemble distributed systems rather than mobile apps. The intelligence lives in a persistent runtime while the phone is one authenticated endpoint among several.

For developers, the mobile app is only part of the job. They also have to build the components that keep an agent running, connect it to a user’s devices, and ensure those connections remain secure.

The agent keeps running independently, while the phone is simply another place to check in, approve actions, or start a conversation.

Looking at OpenClaw alongside Anthropic and OpenAI, it’s hard not to notice the same pattern. The agent keeps running independently, while the phone is simply another place to check in, approve actions or start a conversation. That architecture solves many practical problems, which may explain why several companies are heading in the same direction.

The post OpenClaw’s new app doesn’t run AI on your phone. That’s the whole point. appeared first on The New Stack.

Cloudflare wants to build the economic layer of the AI web

Aerial view of a large multi-level highway interchange with heavy traffic flowing in multiple directions, surrounded by trees, parks and city buildings.

AI has changed the web right before our eyes. With Google’s AI Overviews doing the heavy lifting, publications that once owned the first page of search results are being replaced by summaries. Readers get their answer without ever clicking through. Much of the traffic has simply stopped.

Cloudflare on Wednesday announced a slew of updates for publishers who are facing this new reality. From new crawler classifications and analytics dashboards to Answer Engine Optimization tools and an expansion of its Pay Per Crawl program, it’s clear the company is trying to become the economic pipes of the AI web.

The shift from ‘keep out’ to ‘let’s make a deal’

A year ago, Cloudflare’s pitch was practically defensive, asserting that website owners should be able to block AI crawlers. And while that still holds true, the company has pivoted to discussing building “rails” for an “agentic economy.” And it makes sense. If AI agents are already browsing the web, collecting content, and in some cases buying things, someone needs to handle the business end of how the sites they visit are compensated. Cloudflare thinks that someone should be Cloudflare.

Paying for value, not visits

Roughly a year ago, Cloudflare launched Pay Per Crawl, which let publishers set a price for AI companies to pay when they fetched a page. Now the company is pushing toward Pay Per Use, which means publishers get paid when their content actually appears in an AI-generated answer.

To backtrack, under the old model, an AI crawler pays to visit your site, whether or not it does anything useful with what it finds. Cloudflare says it’s already testing this with Ceramic.ai and You.com, each running slightly different versions of the concept.

Instead of charging for access — which is basically a toll booth — publishers are charging for value.

The economics here flip. Instead of charging for access — which is basically a toll booth — publishers are charging for value. That’s closer to how affiliate marketing or licensing deals work, and it’s a much harder problem to solve. It requires knowing which content contributed to which answer, which means an attribution infrastructure that doesn’t really exist at scale yet.

Credit: Cloudflare.

Crawlers need clearer labels

But here’s where things get more technical — and a bit political.

Cloudflare wants AI companies to stop lumping all their crawlers together. Right now, a single bot from a major AI company might be fetching pages for search indexing, model training, and agent tasks all at once. That makes it impossible for site owners to say yes to one use and no to another.

Starting September 15, Cloudflare plans to change the defaults for new and free-tier sites. AI search crawling stays on, but training and agent access get blocked on ad-supported pages unless the site owner opts in. Mixed-use crawlers that refuse to separate their traffic get blocked entirely.

The company is clearly taking a shot at Google here, noting that Google’s bundled approach gives it access to roughly twice as much content as AI-native competitors. Separating crawlers’ intent has become a prerequisite for any kind of functioning market between publishers and AI companies. Because if site owners can’t distinguish between “index my page for search” and “train your model on my writing,” they’ll increasingly just block everything.

Separating crawlers’ intent has become a prerequisite for any kind of functioning market between publishers and AI companies.

Optimizing for AI answers

Cloudflare is also rolling out a dashboard designed for business teams, called Attribution Business Insights. Think of it as the AI equivalent of knowing your Google Search Console numbers, except the “search engine” is now ChatGPT or Perplexity or whatever agent your reader happened to ask.

On top of that, Cloudflare is introducing Answer Engine Optimization (AEO). The idea is that ranking in Google is no longer enough; for publishers to succeed, they also need to understand how and where their content gets cited in AI-generated responses. That’s a different optimization problem than SEO, and right now almost nobody has good tooling for it.

Infrastructure as competitive advantage

Cloudflare already sits between websites and the internet. It sees the traffic and knows what changed on a page and what didn’t. It can tell a crawler to come back later because nothing’s new, which, by the way, it says would eliminate over 50% of current AI crawl traffic.

That puts Cloudflare in a unique position. It’s already part of the path between AI companies and the web. Now it wants to become the layer that manages access, attribution, and eventually payments between them. If AI agents become the primary way people find information online, Cloudflare is betting the next battle won’t be over who builds the smartest model but instead over who builds the infrastructure everyone else relies on.

If AI agents become the primary way people find information online, Cloudflare is betting the next battle won’t be over who builds the smartest model — it’ll be over who builds the infrastructure everyone else relies on.

The post Cloudflare wants to build the economic layer of the AI web appeared first on The New Stack.

How Cursor deploys AI inside the enterprise

1 July 2026 at 19:03
Pauline Brunet, VP of Forward Deployed Engineering at Cursor, at AIEWF.

Forward deployed engineering has quickly become one of the most prominent roles in enterprise AI. Sitting somewhere between software engineering, product development and customer implementation, forward deployed engineers [FDEs] work directly with organizations to implement AI capabilities.

At Cursor, the role is especially ambitious. Pauline Brunet, the company’s VP of Forward Deployed Engineering, is building a team that works with organizations to implement agents across the entire software development lifecycle.

In an interview with Latent Space at the AI Engineer World’s Fair, Brunet discussed Cursor’s vision of an “AI software factory,” the challenge of expanding agent adoption beyond individual enthusiasts, and what engineers need to demonstrate if they want to move into forward-deployed work.

What forward deployed engineering means at Cursor

Latent Space: To begin with, how do you define forward deployed engineering?

Pauline Brunet: Forward deployed engineering depends on the business, the product, and the customer. You have to consider how configurable the application is. Is it something customers can use out of the box, or are you deploying something complex and highly configurable?

You also have to consider where customers are in their journey.

I don’t think of forward deployed engineering as a team that supports a traditional, out-of-the-box deployment. I think of it as a team that goes on-site, works inside a customer’s systems and tools, and deploys applications or platforms that help solve challenges at scale.

Those deployments are highly configurable and customized around the customer’s workflows, processes, systems, and tools.

Latent Space: Cursor’s customers are predominantly engineers. How does the FDE role apply to the way they use the product?

Brunet: Cursor is an AI coding platform and coding assistant. We work with people on AI-assisted coding, synchronous and asynchronous agents, and ultimately the idea of an AI software factory.

Today, we work with customers across many industries, including financial services, telecommunications, software development, technology, and semiconductors.

We help transformation leaders, IT leaders, and CTO organizations create an AI software factory across their operations. That includes how they plan and design software, how they write code, how they test and review it, and how they deploy and maintain applications at scale. So, very focused on the software development lifecycle from start to finish.

Building Cursor’s FDE team

Latent Space: How large is Cursor’s FDE team?

Brunet: We’re growing rapidly. Our goal is to grow the team tenfold by the end of December.

Latent Space: Are your current FDE employees primarily engineers, or does the team also include product specialists?

Brunet: They are all engineers. We hire software engineers with at least five years of experience and extensive customer-facing experience.

These are people who have developed and shipped code in production. They have built and designed systems, and they can make trade-off decisions and evaluate which systems or technologies should be used.

They also need customer-facing experience. We have people who previously worked at companies including Spotify, Rippling, and Palantir, and who have deployed production systems for customers.

From coding assistants to software factories

Latent Space: You mentioned the term “software factory,” which has begun appearing more frequently in the industry. What does that term mean to Cursor?

Brunet: For Cursor, it is about the software development lifecycle from start to finish: how you plan, design, write, review, test, and deploy code.

Today, those stages are often handled by different teams. You might have a design team, a development team, and a product manager working alongside them. Each group may be optimizing its own work with AI-assisted coding, but the process remains siloed.

We want to help customers across the entire lifecycle. You should be able to say, “Here is the feature I want to develop,” and then have long-running agents work with you across every step. That could include creating the plan and product requirements document, producing a demonstration of what the feature might look like, writing and testing the code, putting it into production, and maintaining it.

Issues and product feedback should also feed back into that same lifecycle. For us, a software factory means long-running agents helping people throughout that entire process.

Latent Space: So it is broader than agent orchestration alone?

Brunet: Correct. Exactly.

Moving beyond individual AI adopters

Latent Space: What problems are enterprises encountering as they try to implement agent technology?

Brunet: One challenge is that adoption is still concentrated among early adopters.

Within an organization, you might have 10% or 20% of people who are enthusiastic early adopters. They have done great work using local agents and cloud agents for their own tasks, and they have become highly productive.

What is missing in the next phase is the ability to use long-running agents across teams, processes, and workflows.

That requires more support from the top of the organization. Leadership has to say, “This is a priority, and this is how we want to automate or change this process.”

For the FDE team, it is therefore important to find the right champions inside an organization: people who want to meaningfully change the business and who will work with us and their internal teams to transform how work gets done.

Standardizing work with cloud agents

Latent Space: Local AI appears to be gaining momentum, partly because of the increasing availability of open-source models. Are you doing more local AI implementation work with customers?

Brunet: We have local agents that people run through the desktop application or the CLI, and that experience is largely self-service. People have adopted the technology at a phenomenal rate, particularly across Cursor’s user base.

We are also seeing people adopt cloud agents because they are excited about being able to run tasks without keeping their laptops half open. Agents can now work in the cloud on tasks that previously ran locally.

What becomes interesting is when this moves beyond an agent helping with one person’s job. The next question is how agents can work across a function, team, or organization so that processes are automated consistently. For example, you could have a QA agent applying the same process across several development teams.

We are receiving a lot of questions from customers about those kinds of use cases.

How customer deployments influence Cursor’s roadmap

Latent Space: Do the lessons from these deployments feed back into the core Cursor product?

Brunet: Yes. The forward deployed engineering team works very closely with customers on their use cases, so we are naturally a good way for the product and engineering teams to understand what customers want to build next.

We work closely with those teams and play a significant role in helping shape Cursor’s product roadmap.

The changing role of the forward deployed engineer

Latent Space: As agents become more autonomous, how do you expect the FDE role to evolve?

Brunet: I think the role is going to change drastically. I always say that if we are doing the same job we were doing six months ago, we have done something wrong.

Right now, people are still looking for inspiration about the use cases they can solve, so we want to propose new possibilities.

In software development, for example, we can show how designers and product managers might work seamlessly in Cursor alongside developers and testing teams.

We might also ask whether a company has considered using long-running agents to handle call-center or ticketing processes from start to finish.

As we work across industries such as healthcare, life sciences, the public sector, retail, and consumer packaged goods, we will continue identifying use cases across marketing, sales, and supply-chain operations. The FDE role will evolve alongside those possibilities.

How engineers can prepare for an FDE career

Latent Space: There are around 7,000 AI engineers at this conference. What advice would you give developers who want to move into forward deployed engineering?

Brunet: I’ve had this conversation five or six times already today. We are looking for builders with software engineering experience: people who have identified a problem and built a production-grade application or system from start to finish.

You should have designed it, developed it, tested it, and put it into production with real users.

My recommendation is to find those kinds of projects inside your organization and take ownership of them from beginning to end. Make sure you understand why you made each design decision.

How did you select the database? How did you choose the different services? Why did you design the system in that particular way? What were the trade-offs?

You should also understand the measurable return on investment, both in traditional business terms and through evaluations that demonstrate the value you are creating for internal customers.

If you want to get into forward deployed engineering, become familiar with these kinds of projects, gain experience delivering them, and learn how to explain the decisions you made.

2026 BAIR Graduate Showcase

Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more. Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for the better.

Now they are headed everywhere ideas travel: to faculty and postdoctoral positions, to industry research labs, and to startups of their own founding — and several are still exploring what comes next and would love to hear from you.

Please join us in celebrating the achievements of these wonderful graduates. We are proud of everything they have accomplished at Berkeley, and we can’t wait to see what they do next!

Thank you to our friends at the Stanford AI Lab for this idea!


Baifeng Shi

Baifeng Shi


Email: baifeng_shi@berkeley.edu
Website: https://bfshi.github.io/
Advisor(s): Trevor Darrell
Research Blurb: I work on building generalist vision and robotic models.
What's next: Member of Technical Staff at Physical Intelligence

Charlie Snell

Charlie Snell


Email: csnell22@berkeley.edu
Website: https://sea-snell.github.io
Advisor(s): Dan Klein
Research Blurb: My work aims to understand when and how the different LLM scaling paradigms can be traded off and interchanged. In particular, test-time scaling treats each prompt independently, drawing long chains of inferences and then forgetting them entirely between prompts. This differs critically from pretraining, which instead learns a compressed representation from a large dataset. I believe bridging the gap between these methods of scaling computation, presents a key open challenge in the field: how can we develop methods which turn the inferences drawn at test-time back into learned representations that the model can hold onto across interactions.

Devin Guillory

Devin Guillory


Email: dguillory@berkeley.edu
Website: https://devinguillory.com
Advisor(s): Trevor Darrell
Research Blurb: Accounting for data shifts in computer vision models
What's next: Building collaborative AI systems, looking for conspirators.

Eve Fleisig

Eve Fleisig


Email: efleisig@berkeley.edu
Website: https://efleisig.com
Advisor(s): Dan Klein
Research Blurb: I design language models to work reliably and fairly for the broad range of real LLM users. First, my research leverages disagreement among user preferences as signal, in order to train and evaluate LLMs for entire populations of users. Second, I work on designing rigorous evaluations to extricate challenging LLM harms that diverse users face. Finally, I work on core technical failures of LLMs, like miscalibrated confidence, to reduce downstream risks when models are deployed to users with different needs. Combined, these interventions facilitate building LLMs that minimize societal harms, and maximize benefits to a wider range of real-world users.
What's next: Postdoctoral fellow at Princeton CITP

Grace Luo

Grace Luo


Email: graceluo@berkeley.edu
Website: https://graceluo.net
Advisor(s): Trevor Darrell
Research Blurb: My research is on interpreting and controlling generative models. For example, I've worked on re-purposing image generators for computer vision tasks, and meta-modeling language activations for better LLM probing and steering.
What's next: Research scientist in industry

Hanlin Zhu

Hanlin Zhu


Email: hanlinzhu@berkeley.edu
Website: https://hanlinzhu.com/
Advisor(s): Stuart Russell, Jiantao Jiao
Research Blurb: My research centers on understanding and improving the reasoning capabilities of large language models (LLMs).
What's next: Member of Technical Staff at OpenAI

Haozhi Qi

Haozhi Qi


Email: hqi@berkeley.edu
Website: https://haozhi.io/
Advisor(s): Jitendra Malik, Yi Ma
Research Blurb: Dexterous Manipulation and Robot Learning
What's next: Research scientist at Amazon; Faculty at University of Chicago

J.D. Zamfirescu-Pereira

J.D. Zamfirescu-Pereira


Email: zamfi@berkeley.edu
Website: https://zamfi.net
Advisor(s): Bjoern Hartmann
Research Blurb: My research focuses on effective human-AI co-design. I study the boundaries of language interfaces as a medium for interacting with AI, creating systems that blend language-focused interactions with structured user interfaces that draw on different levels of abstraction. I focus on language-oriented technologies, like LLMs and text-to-image models, that are powerful mediators of design processes. These technologies enable humans to describe their desires at almost any level of abstraction, from high-level goals vaguely specified (“I’d like a game to help my kid learn to read”) to low-level corrections of undesired outputs (“Don’t say ‘I know because I’ve tasted it’ when about a recipe substitution's taste”).
What's next: Assistant Professor, Computer Science, UCLA

Jiachen Lian

Jiachen Lian


Email: jiachenlian@berkeley.edu
Website: https://jlian2.github.io
Advisor(s): Gopala Anumanchipalli
Research Blurb: My research focuses on human-centered AI across speech, healthcare, and systems.
Looking for: Look for AI talents to join our startup

Josh Kang

Josh Kang


Email: minwoo_kang@berkeley.edu
Website: https://joshuaminwookang.github.io/
Advisor(s): John Canny
Research Blurb: I study language modeling and related topics in NLP; specific interests are human user simulation and building conversational, collaborative AI agents.
What's next: AI Scientist at Mistral AI

Junhao (Bear) Xiong

Junhao (Bear) Xiong


Email: junhao_xiong@berkeley.edu
Website: https://www.linkedin.com/in/junhao-bear-xiong
Advisor(s): Jennifer Listgarten, Yun Song
Research Blurb: Junhao (Bear) Xiong is a PhD candidate at UC Berkeley, advised by Jennifer Listgarten and Yun S. Song. His work focuses on machine learning methods for biology, with an emphasis on generative modeling for proteins. Previously, he studied Applied Math and Computer Science at Johns Hopkins.
Looking for: Research scientist

Kaylo Littlejohn

Kaylo Littlejohn


Email: kaylo_littlejohn@berkeley.edu
Website: https://kaylolittlejohn.com
Advisor(s): Gopala Anumanchipalli
Research Blurb: My research is focused on speech modeling and natural language processing. I co-led the development of multimodal AI tools to accurately translate brain activity into text, audible personalized speech, and a high-fidelity "digital talking avatar" (Nature 2023, Nature Neuroscience 2025). I am also tech lead for voice modeling at Roblox.
Looking for: Research Scientist / Engineer

Kent Chang

Kent Chang


Email: kentkchang@berkeley.edu
Website: https://kentkc.org
Advisor(s): David Bamman
Research Blurb: I work on NLP and multimodal machine learning, with a focus on evaluating large language models and building multimodal systems for understanding dialogue, narrative, and social interaction. My research includes benchmarks for LLM memorization, multimodal datasets sourced from feature films and television, and studies of model behavior. I'm interested in bridging computational methods with questions from the humanities and social sciences about whose voices get represented in AI systems, and about AI's broader impact. My work has appeared at EMNLP and ACL, among others.
Looking for: (teaching) faculty, Research Scientist, ML/AI SWE

Kevin Black

Kevin Black


Email: kvablack@berkeley.edu
Website: https://kevin.black
Advisor(s): Sergey Levine
Research Blurb: I work on large-scale robot learning: including imitation learning, reinforcement learning, generative modeling, real-time control, and whatever else it takes to make robots work in the real world!
What's next: Research Scientist of Physical Intelligence

Kunhe Yang

Kunhe Yang


Email: kunheyang@berkeley.edu
Website: https://www.kunheyang.com/
Advisor(s): Nika Haghtalab
Research Blurb: My research focuses on the theoretical foundations of designing and evaluating AI algorithms in environments shaped by human incentives and AI agency. My work spans human-centric policy learning, incentive-aware evaluation, and multi-agent collaboration and information transmission, drawing on tools from machine learning theory and computational economics.
What's next: Postdoc Research at Stanford

Lisa Dunlap

Lisa Dunlap


Email: lisabdunlap@berkeley.edu
Website: https://lisabdunlap.com
Advisor(s): Joseph Gonzalez, Trevor Darrell
Research Blurb: Auditing generative models.
What's next: Research Engineer at Anthropic

Long (Tony) Lian

Long (Tony) Lian


Email: longlian@berkeley.edu
Website: https://tonylian.com/
Advisor(s): Trevor Darrell, Adam Yala
Research Blurb: My research primarily focuses on developing real-time multi-modal multi-agent systems and parallel reasoning systems through end-to-end RL.
What's next: Member of Technical Staff at Thinking Machines Lab

Maulik Bhatt

Maulik Bhatt


Email: maulikbhatt@berkeley.edu
Website: https://maulikb.com
Advisor(s): Negar Mehr
Research Blurb: My research develops autonomous robots that can safely coordinate with humans and other robots in shared environments. I build scalable algorithms grounded in game theory and diffusion models that let agents reason about the intent and behavior of others around them. My work spans real-time multi-agent trajectory planning and imitation learning in the presence of multi-modality. I've validated these methods on hardware platforms ranging from quadrotors to manipulators, with the goal of making multi-agent coordination robust, interpretable, and deployable in the real world.
What's next: Joining Toyota Woven's end-to-end autonomous driving team.

Michael Psenka

Michael Psenka


Email: psenka@berkeley.edu
Website: https://www.michaelpsenka.io/
Advisor(s): Aditi Krishnapriyan
Research Blurb: Work in various domains (reinforcement learning, world models, AI+bio/chem), generally working on longer-horizon and out-of-distribution problems in planning and interpolation (e.g. robot manipulation from start state to goal, molecular dynamics of proteins between ground states). My thesis took a variational approach (think calculus of variations) directly from deep generative models of the environment, framing path-finding as minimizing a functional induced by the learned model itself (its score, its critic, or its dynamics). Through my research I've gained insight on how to properly handle dynamics in deep learning systems, and I plan to continue developing systems that are dynamic and adaptive.
What's next: Lead Research Scientist at Baseten

Nathan Lichtlé

Nathan Lichtlé


Email: nathan.lichtle@gmail.com
Website: https://nathanlichtle.com
Advisor(s): Alexandre M. Bayen
Research Blurb: RL for autonomous driving.
What's next: Chief Scientist & Co-founder at Yumi Health

Neerja Thakkar

Neerja Thakkar


Email: nthakkar@berkeley.edu
Website: https://neerja.me/
Advisor(s): Jitendra Malik
Research Blurb: My research focuses on scaling predictive world models to handle the complexity of in-the-wild motion. Using autoregressive and diffusion frameworks, I develop better representations for real-world prediction and propose methods to efficiently adapt these models to new domains.
Looking for: Research scientist

Nikita Mehandru

Nikita Mehandru


Email: nmehandru@berkeley.edu
Website: https://n-mehandru.github.io/
Advisor(s): Ahmed Alaa and David Bamman
Research Blurb: My research develops and applies machine learning methods for clinical reasoning and disease progression modeling using unstructured text and time series data from electronic health records. In collaboration with physicians at UCSF, I bridge method development and clinical validation with the intention to build reliable, interpretable AI systems in medicine.
Looking for: Research Scientist

Niklas Lauffer

Niklas Lauffer


Email: nlauffer@berkeley.edu
Website: https://niklaslauffer.github.io/
Advisor(s): Stuart Russell and Sanjit Seshia
Research Blurb: Niklas's research is focused on AI safety and reinforcement learning, particularly in the area of multi-agent interaction and LM agents. He's worked on enabling adversarial learning in cooperative and mixed-motive settings, solving issues of covariate shift in training LM agents on long-horizon tasks, as well as evaluating safety risks posed by LM agents in multi-agent settings.
What's next: Research Scientist at Google Deepmind

Qiyang Li

Qiyang Li


Email: qcli@berkeley.edu
Website: https://colinqiyangli.github.io/
Advisor(s): Sergey Levine
Research Blurb: Recent progress in robotic manipulation policy learning has been largely driven by (1) the increasing availability of large-scale prior datasets and (2) the success of action chunking, where the policy predicts a short sequence of future actions rather than a single one. However, most action chunking policies are trained via supervised imitation learning, because efficient online self-improvement with reinforcement learning (RL) remains challenging—limiting real-world applicability. My PhD research studied how we could leverage prior data to optimize action-chunking policies with RL, combining empirical results with theoretical insights.
Looking for: Post-doc/research scientist for RL in robotics and LLMs!

Sampada Deglurkar

Sampada Deglurkar


Email: sampada_deglurkar@berkeley.edu
Website: https://sdeglurkar.github.io/
Advisor(s): Prof Claire Tomlin
Research Blurb: My research is in providing safety assurances for AI-enabled autonomous systems, ranging from robots to autonomous vehicles to aviation systems. For this, I have worked with uncertainty quantification for machine learning models, decision-making under uncertainty algorithms, and tools for producing probabilistic guarantees on system operation.
Looking for: Research scientist, Research engineer

Vinamra Benara

Vinamra Benara


Email: vbenara@berkeley.edu
Website: https://cs.berkeley.edu/~vbenara
Advisor(s): Ion Stoica
Research Blurb: My research focuses on LLM post-training, including data curation, RLHF, RLVR with VLMs, evaluations, reasoning, agentic workflows, and interpretability. I also have strong expertise in systems infrastructure for distributed computing.
Looking for: Research scientist / Research Engineer

Vongani Maluleke

Vongani Maluleke


Email: vongani_maluleke@berkeley.edu
Website: https://people.eecs.berkeley.edu/~vongani_maluleke/
Advisor(s): Jitendra Malik and Angjoo Kanazawa
Research Blurb: Vongani Maluleke is a PhD candidate at UC Berkeley (BAIR, advised by Jitendra Malik and Angjoo Kanazawa), where she led the development of MAGNet, a unified multi-agent motion generation framework that supports a wide range of motion generation tasks without retraining or architectural changes, outperforming task-specialized state-of-the-art baselines. She is currently extending this work by deploying it on a Unitree G1 humanoid to make it embody social intelligence. Before her PhD, she was a Senior AI Consultant at Deloitte, awarded Exceptional Performer two consecutive years, leading AI system development across media, telecommunications, retail, and financial services.
Looking for: Research scientist

Wei-Jer Chang

Wei-Jer Chang


Email: weijer_chang@berkeley.edu
Website: https://weijer-chang.github.io/
Advisor(s): Masayoshi Tomizuka
Research Blurb: My research focuses on developing safe and intelligent autonomous systems for complex, human-centered environments. I work at the intersection of machine learning, generative models, and reinforcement learning, with applications in autonomy. My work addresses challenges in multi-agent interaction, interactive human behavior, and long-tail safety-critical scenarios at scale.
Looking for: Research Scientist, Applied Scientist, Roboticist

Xiuyu Li

Xiuyu Li


Email: xiuyu@berkeley.edu
Website: https://xiuyuli.com/
Advisor(s): Kurt Keutzer
Research Blurb: My research focuses on developing scalable and self-improving large language model agents, with emphasis on coding agents for complex, long-horizon tasks. This direction builds on my work in parallel reasoning, and on broader expertise in making generative models more efficient in training and inference across language and vision.
What's next: Member of Technical Staff at xAI

Yichen Xie

Yichen Xie


Email: yichenxie0928@gmail.com
Website: https://yichen928.github.io/
Advisor(s): Masayoshi Tomizuka
Research Blurb: My research focuses on building multimodal foundation models and world models that understand and interact with complex physical environments. I aim to develop unified representations across modalities, enabling AI systems to reason over space, time, and dynamics toward general-purpose embodied intelligence.
What's next: Research Scientist at Luma AI

Yigit Efe Erginbas

Yigit Efe Erginbas


Email: erginbas@berkeley.edu
Website: https://www.linkedin.com/in/erginbas/
Advisor(s): Kannan Ramchandran, Thomas A. Courtade
Research Blurb: My PhD research spans two threads: online learning in large-scale markets, and interpretability of large machine learning models. In the first, I work on sequential decision-making with applications to recommendation, pricing, and assortment selection. My focus is on designing algorithms with provable guarantees for welfare maximization, revenue maximization, and stability. In the second, I develop scalable attribution methods that exploit the sparse, low-degree structure of real-world interactions, using tools from signal processing and information theory. More recently, I have been exploring principled ways to evaluate the faithfulness of model self-explanations.
What's next: Researcher at Hudson River Trading's AI Labs (HAIL)

Yiheng Li

Yiheng Li


Email: yhli@berkeley.edu
Website: https://Yihengli.com
Advisor(s): Masayoshi Tomizuka
Research Blurb: I am working on vision world modeling, with prior experience in diffusion model's efficiency as well as in autonomous driving.
What's next: Research Scientist at Waymo

Zhe Fu

Zhe Fu


Email: zhefu@berkeley.edu
Website: https://fu-zhe.com/
Advisor(s): Alexandre Bayen
Research Blurb: My research focuses on physics-informed learning and control for mixed-autonomy systems, with applications in transportation. I design physics-informed neural networks to learn solutions of nonlinear partial differential equations, enabling accurate and data-efficient prediction of traffic dynamics. Building on these models, I develop both model-based and learning-based control strategies that coordinate automated vehicles to improve system-level performance. My work bridges machine learning, control, and real-world deployment, and has been validated in large-scale field experiments. More broadly, I aim to advance trustworthy, interpretable AI for decision-making in complex, real-world systems.
What's next: I will be an Energy Fellow at Stanford after graduation. Also looking for Faculty, or research scientist positions in AI, control, and autonomy.

ML Development in VS Code with Google Cloud Power: Workbench Extension Now Available

1 July 2026 at 17:15
The Google Cloud Workbench Notebooks extension for VS Code has officially launched, allowing developers to connect their local IDE to scalable, cloud-based Jupyter environments. This integration streamlines the machine learning lifecycle by eliminating context switching and providing direct access to high-performance Google Cloud infrastructure. To support transparency and community-driven innovation, the newly released extension is fully open-sourced and available on GitHub and the VS Code Marketplace.

Why we built ADK 2.0

1 July 2026 at 17:15
Answering the questions of "why we built ADK 2.0". This explains the rationale, some of the features, and why a developer should consider upgrading. This will be published the day after ADK go 2.0 launches.

Mastering Agentic Techniques: AI Agent Reinforcement Learning

1 July 2026 at 17:04
Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer...

Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer reinforcement learning with verifiable rewards (RLVR) workflows for reasoning and agent tasks. RL is now becoming a practical technique for specialized AI where enterprises need more accurate agents for domain-specific workflows.

Source

How Amazon tracks carbon intensity across its operations

1 July 2026 at 15:56
At Amazon, we believe that as our strategy to reach net-zero by 2040 evolves, we need to continue to raise the bar on what and how we measure. Measuring carbon emissions across our entire business is complex, and the tools and methodologies available are continuously improving. Each year, we report on our carbon intensity, absolute emissions, and sustainability progress in our Sustainability Report, and we continue to adopt more-precise tools and methodologies to ensure our data is as meaningful and as representative of our decarbonization journey as possible. Carbon intensity — the amount of emissions per unit of activity or production — is one of the most important tools for tracking decarbonization progress. But not all activities are the same. Across industries, some of the most meaningful intensity metrics are sector specific, tied to the actual activity being decarbonized, across buildings, energy, transportation, and products and beyond. Examples of metrics used by companies include carbon dioxide equivalent per megawatt-hour for electricity, per square foot for building decarbonization, and per kilometer for transport. At Amazon, we apply the same principle: we are developing intensity metrics tailored to specific business activities, so we can measure what matters most. Emissions per unit shipped For Amazon's retail operations, we track carbon emissions per unit shipped. This metric matters because it directly reflects the efficiency of our delivery operations at scale. We’ve reduced emissions per unit shipped every year since 2019, resulting in a carbon intensity reduction of 39% at the end of 2025, relative to 2019. As we ship more packages, we're doing so with less carbon per unit, and our investments in carbon-free energy, smarter routing, lighter packaging, low-carbon fuels, alternative transportation methods (like rail versus road or air), and electric vehicles are all translating into real per-unit reductions. We also track carbon intensity at the regional and country level to understand geographic variation and target interventions accordingly. It's the clearest measure of whether we're decoupling delivery growth from emissions growth. Why this is unique to Amazon Amazon is more than a traditional logistics company: we're also a pharmacy, a grocery store, a cloud services company, a movie studio, a device manufacturer, a satellite business, and much more. As a result, our emissions span the entire value chain — from the manufacture of goods, long-haul shipping, and air freight to warehousing and distribution and middle-mile and last-mile delivery. Amazon’s decarbonization efforts require strategies that can address significant portions of global transportation and supply chains simultaneously. That breadth creates complexity — and opportunity. Solutions that we develop together with our suppliers and partners don’t only stay within the four walls of Amazon; they have the potential to drive decarbonization across multiple industries at once. This is a responsibility we take seriously: we’re proud of our progress and want to share what we are learning along the way. Simplifying our economic intensity metric Part of sharing what we learn is being transparent about how we measure and what we change. One change that we’ve made recently is in our economic carbon intensity indicator. In our 2025 reporting, we updated this from grams of carbon dioxide equivalent per dollar of gross merchandise sales (gCO2e/$GMS) to grams of carbon dioxide equivalent per dollar of revenue (gCO2e/$Revenue). Revenue is publicly reported, widely understood, and aligns with how most companies disclose the carbon intensity of their operations— making our progress easier to compare with the rest of the industry’s. This update does not change our year-over-year trajectory. Looking ahead Climate science and carbon accounting are not static, and as the science and our own operations evolve, we'll continue to refine how we track and report our progress. We’re proud of our Climate Pledge goal to reach net-zero carbon by 2040. We’ll continue updating how we measure, so our data always reflects the most precise and meaningful picture of where we are and where we’re headed. To learn more about Amazon's carbon methodology, visit our Sustainability reporting website.

Build agentic full-stack apps with Genkit

1 July 2026 at 16:01
The open-source Genkit framework has introduced the Agents API, a full-stack tool designed to simplify the complex plumbing of conversational AI by packaging message history, tool loops, and streaming into a single interface. The API supports flexible, server- or client-managed state persistence—allowing for advanced workflows like history branching, long-running detached tasks, and multi-agent coordination—while seamlessly connecting backends to frontends via a unified wire protocol. Currently available in preview for TypeScript and Go, it also integrates with the Genkit Developer UI to allow developers to easily test, debug, and inspect agent snapshots without writing client code.

Anthropic is bringing back Claude Fable 5 globally after US lifts export control order — where can enterprises access it?

Anthropic is restoring global access to its most powerful generally released AI model yet, Claude Fable 5, today, after the U.S. Department of Commerce last night withdrew the emergency export controls it had issued previously around the model.

The U.S. export control order issued on June 12, 2026, led Anthropic to suspend all global access to both Fable 5 and its less restricted cybersecurity counterpart model Claude Mythos 5, just days after both models were initially introduced.

Now, Fable 5 is once again being made available for users globally across the primary Anthropic ecosystem, including the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. The official Claude account on X announced the return of the model at 3:31 pm ET on July 1, 2026.

For organizations leveraging cloud hyperscalers, Anthropic says it is moving to re-enable access on Amazon Web Services, Google Cloud, and Microsoft Foundry “as quickly as possible.” So far, VentureBeat's research has been unable to confirm if the models have been restored on these external cloud hyperscaler platforms yet.

Mythos 5 remains a different case. A letter posted on the social network X allegedly from U.S. Commerce Secretary Howard Lutnick to Anthropic executive Tom Brown says a license is no longer required for the export, reexport, or in-country transfer of Fable and Mythos.

But Anthropic’s own redeployment post on its website says only that Mythos 5 access has been restored for “a set of US organizations,” following government approval on June 26. The company says it is continuing to coordinate with the government to expand access to broader domestic and international partners in its opt-in cybersecurity testing program, Project Glasswing.

That leaves Mythos 5 in a middle category: legally cleared from the emergency export-control order, but not generally available. The current limit appears to come from Anthropic’s decision to keep Mythos behind a vetted-access model, with the U.S. government still playing a role in approvals, standards and expansion.

Posting on X, Commerce Secretary Howard Lutnick said Anthropic and the government had “worked closely” to “analyze and approve Fable 5,” while White House Chief of Staff Susie Wiles also posted on X, framing the decision around U.S. AI leadership and deployment speed.

Wiles wrote that the United States is the “undisputed winner in the AI race,” adding that the shared priority is to “get the best tech deployed as quickly and safely as possible.”

The reversal follows concerns from cybersecurity leaders and AI policy experts over the export control order, who argued that the U.S. risked hobbling its own industry while giving Chinese AI labs an opening. Former Facebook security chief Alex Stamos called the Fable restriction a “huge own goal for the US,” warning that security companies could be driven toward Chinese models, while other critics said the so-called "ad hoc" regulatory intervention made dependence on U.S. AI platforms look like a strategic liability.

Reminder on Claude Fable 5 pricing

For chief information and technology officers evaluating the return of the model, the deployment comes with distinct structural conditions and significant financial investments.

Anthropic is pricing both Fable 5 and Mythos 5 at $10.00 per million input tokens and $50.00 per million output tokens, the most expensive of all frontier models globally.

Model

Input ($/1M)

Output ($/1M)

Total ($/1M)

Source

MiMo-V2.5 Flash

$0.10

$0.30

$0.40

Xiaomi

deepseek-v4-flash

$0.14

$0.28

$0.42

DeepSeek

deepseek-v4-pro

$0.435

$0.87

$1.305

DeepSeek

MiniMax-M3

$0.30

$1.20

$1.50

MiniMax

LongCat-2.0 — limited-time promo

$0.30

$1.20

$1.50

LongCat

Gemini 3.1 Flash-Lite

$0.25

$1.50

$1.75

Google

Qwen3.7-Plus

$0.40

$1.60

$2.00

Alibaba Cloud

MiMo-V2.5

$0.40

$2.00

$2.40

Xiaomi

LongCat-2.0 — standard

$0.75

$2.95

$3.70

LongCat

Grok 4.3 (low context)

$1.25

$2.50

$3.75

xAI

MiMo-V2.5 Pro (≤256K)

$1.00

$3.00

$4.00

Xiaomi

Kimi-K2.6

$0.95

$4.00

$4.95

Moonshot AI

GLM-5.2

$1.40

$4.40

$5.80

Z.ai

GPT-5.6 Luna

$1.00

$6.00

$7.00

OpenAI

Grok 4.3 (high context)

$2.50

$5.00

$7.50

xAI

MiMo-V2.5 Pro (>256K)

$2.00

$6.00

$8.00

Xiaomi

Qwen3.7-Max

$2.50

$7.50

$10.00

Alibaba Cloud

Gemini 3.5 Flash

$1.50

$9.00

$10.50

Google

Gemini 3.1 Pro Preview (≤200K)

$2.00

$12.00

$14.00

Google

GPT-5.6 Terra

$2.50

$15.00

$17.50

OpenAI

GPT-5.4

$2.50

$15.00

$17.50

OpenAI

Gemini 3.1 Pro Preview (>200K)

$4.00

$18.00

$22.00

Google

Claude Opus 4.8

$5.00

$25.00

$30.00

Anthropic

GPT-5.5

$5.00

$30.00

$35.00

OpenAI

GPT-5.5 Instant (chat-latest)

$5.00

$30.00

$35.00

OpenAI

Sakana Fugu Ultra (≤272K)

$5.00

$30.00

$35.00

Sakana AI

GPT-5.6 Sol

$5.00

$30.00

$35.00

OpenAI

Claude Fable 5 / Claude Mythos 5

$10.00

$50.00

$60.00

Anthropic

However, to incentivize immediate enterprise adoption following the export control order disruption saga, Anthropic is executing a temporary rollout plan through July 7.

For Pro, Max, Team, and select Enterprise subscriptions, Fable 5 usage will be included at no added cost for up to 50% of a user’s weekly tier allowance.

After July 7, Fable 5 will move to usage credits for those plans. For standard Enterprise seats, there is no included Fable 5 allowance; all usage is billed through credits, and the model will not work for those users unless credits are enabled.

Already, some AI influencers are attempting to offer enterprises and developers guidance on how to maximize their usage of Fable 5 during its 7-day discounted price/subscription included promotion:

Chronology of a Crisis: From Launch to Lockout

The whiplash regulatory cycle surrounding the model underscores the volatility currently facing enterprise software supply chains. The crisis unfolded over a rapid, three-week timeline:

  • June 9, 2026: Anthropic launches Claude Fable 5 and Mythos 5. Early corporate case studies report major performance gains. For instance, Stripe reports that Fable 5 compressed a codebase-wide migration across a 50-million-line Ruby infrastructure into a single day — a project estimated to take a team more than two months by hand.

  • June 12, 2026: At 5:21 PM ET, the U.S. government issues an export-control directive citing national security authorities. The order bans access to the models by any foreign national, whether inside or outside the borders of the United States. Lacking real-time mechanisms to verify user nationality at the API layer, Anthropic is forced to pull the plug for all customers to ensure compliance. Anthropic says access to all other Anthropic models was not affected.

  • June 13–25, 2026: Enterprise users and developers face abrupt disruption, forcing workflows that had adopted Fable 5 or Mythos 5 to fall back to older models such as Opus 4.8. Tensions peak as Anthropic publicly objects, arguing that pulling a major commercial model over a narrow jailbreak finding could “essentially halt all new model deployments for all frontier model providers.”

  • June 26, 2026: The U.S. government allows Anthropic to restore Mythos 5 access to a set of trusted U.S. organizations, partially reversing the June 12 order. Anthropic says it is restoring access for those organizations and continuing to work with the government to expand Mythos 5 access and make Fable 5 generally available again.

  • June 30, 2026: Commerce Secretary Howard Lutnick sends a letter withdrawing the June 12 export-control license requirement for both Mythos and Fable. The decision removes the emergency legal block, but Anthropic’s rollout still treats the models differently: Fable 5 returns globally, while Mythos 5 remains limited to approved users through Glasswing and related trusted-access channels.

The Technical Catalyst: The Amazon Vulnerability Report

The swift intervention by the federal government stemmed from a report by Amazon researchers describing a method for bypassing Fable 5’s safeguards. This was a brutal irony for Anthropic, given Amazon was one of the startup's initial and largest backers to the tune of $8 billion, and the two companies previously collaborated on improving Amazon's Alexa+ voice assistant.

According to Anthropic, the technique prompted Fable 5 to identify software vulnerabilities; in one case, the model produced code demonstrating how the relevant vulnerability could be exploited.

When the report reached government officials, it triggered alarm regarding the offensive cyber capabilities of public, AI large language models (LLMs). Anthropic countered that the exploit did not tap into unique “Mythos-level” cyber capabilities, noting that its own testing found other models — including Claude Opus 4.8, OpenAI’s GPT-5.5, and Moonshot’s Kimi K2.7 — could identify the same vulnerabilities. Anthropic also said every model it tested could produce the same exploit demonstration as Fable 5.

To break the regulatory logjam, Anthropic developed an improved automated safety classifier specifically trained to catch and neutralize the Amazon technique. Tested by the Commerce Department’s Center for AI Standards and Innovation (CAISI), the updated classifier successfully halts that specific technique in more than 99% of cases.

Anthropic explicitly warns enterprise clients that this safety enforcement comes at an operational cost. Because the new classifiers require an expanded “safety margin” to catch ambiguous edge cases, benign coding and debugging requests may be flagged more often. When a prompt is blocked by the safety layer, the active session automatically downgrades, routing the request to Opus 4.8.

In a post on X, Thariq Shihipar, a Member of Technical Staff at Anthropic working on Claude Code, said that Anthropic is “continuing to refine these safeguards to better distinguish genuine misuse from legitimate requests and reduce false positives.”

Backroom Diplomacy: The Shifting of the Guard

The breakthrough that brought Fable 5 back to commercial markets was as much political as it was technical. According to WIRED, Anthropic initially argued that the administration’s security concerns were overblown and that no frontier model provider could guarantee zero jailbreaks.

That argument frustrated the administration, according to WIRED’s reporting. In recent weeks, Anthropic changed tack, focusing less on the theoretical impossibility of eliminating jailbreaks and more on building stronger safeguards and satisfying the government’s operational concerns.

WIRED reported that Anthropic CEO Dario Amodei was recently replaced in meetings by Brown, whom officials liked more personally. Brown is also the addressee of Lutnick’s June 30 Commerce letter.

Under Brown’s guidance, Anthropic appears to have moved from arguing over the absolute limits of model safety to committing to the expanded safeguards and collaboration framework the administration demanded.

The resulting Commerce letter describes several commitments by Anthropic. Under the terms of the clearance, Anthropic has agreed to:

  1. Proactively detect and address security risks associated with the models.

  2. Work with the U.S. government on protocols, standards and releases for Mythos, Fable and future models.

  3. Inform the U.S. government of malicious activity.

Separately, Anthropic says it will expand pre-release government access and evaluation for frontier models, share information rapidly when significant jailbreaks or misuse patterns are identified, dedicate resources to joint government research and work toward a common industry security bar.

The U.S. Commerce Department explicitly reserved the right to re-evaluate these permissions and re-impose license requirements if circumstances change or if Anthropic fails to meet its commitments.

The Sovereign Calculus: Lessons for Enterprise AI

The two-week blackout of Claude Fable 5 exposed the fragility of centralized, closed-API models for modern business infrastructure. It showed that enterprise automation pipelines remain vulnerable to sudden regulatory shifts and vendor compliance mandates.

The tech community’s response highlights a broader push toward hardware and model sovereignty. Following the initial shutdown, prominent tech figures voiced concerns over this centralization. AI founder Alex Finn described the Anthropic freeze as a major “wakeup call,” urging developers to invest heavily in local, open-weights infrastructure to insulate operations from federal volatility. As Finn noted on social media:

“No company or government will EVER be able to take away your local models.”

For enterprise architects, the return of Fable 5 demands a balanced approach to deployment:

  • The Frontier Performance Advantage: Utilizing closed models like Fable 5 offers state-of-the-art capabilities across agentic coding, long-context work, document reasoning and multi-step enterprise automation, according to Anthropic’s launch materials and early customer examples.

  • The Mitigating Data Trade-Off: Accessing Fable 5 means accepting Anthropic’s mandatory 30-day data retention requirement for covered models. Anthropic says prompts and model completions are retained for at least 30 days by default and then automatically deleted, except when they are part of a safety investigation or must be kept for legal reasons. Highly regulated financial, healthcare and legal groups must evaluate whether this telemetry window complies with their data privacy mandates.

The truth is, enterprises in the U.S. and globally have more options than ever for frontier-class LLMs, especially with the recent launch over the last few months of new, powerful, open weights Chinese alternatives that can be downloaded, run locally or on virtual private clouds, and customized to any enterprise's liking.

MiniMax M3 pairs frontier-tier coding and agentic performance with a 1 million-token context window and native multimodality. Z.ai’s GLM-5.2's benchmark results exceed OpenAI's GPT-5.5 on SWE-bench Pro and several long-horizon coding tests, and near Claude Opus 4.8 on FrontierSWE and MCP-Atlas. Meituan’s LongCat-2.0 is also positioned around enterprise use, with a 1 million-token context window, MIT licensing and strong early developer traction through its Owl Alpha run on OpenRouter — though as we reported, the full weights are still listed as “coming soon.”

Meanwhile, Anthropic's top domestic rival OpenAI is still struggling to release its latest models broadly due to U.S. government pressure. The company says its newest and most powerful models, GPT-5.6 Sol, Terra and Luna — unveiled last week — are starting in a limited preview for a small group of trusted partners after OpenAI previewed the models and their capabilities to the U.S. government and the government requested the rollout be staggered.

OpenAI says it still plans broader availability, but argued in its announcement that "we don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them. We are taking this short-term step because we believe it is the strongest path to broader availability in the coming weeks, while we work with the Administration to develop the cyber Executive Order framework and a repeatable process for future model releases."

The executive order in question, signed by President Donald J. Trump on June 2, 2026, calls upon various federal agencies to collaborate on a process for benchmarking and assessing capabilities of new AI models to ensure they are safe and appropriate for wide release, a process supposed to take 30 days (which would seem to indicate the agencies are due to provide their process tomorrow, July 2, 2026.)

Frontier model launches are starting to look less like ordinary product releases and more like negotiated deployments shaped by U.S. national security review — a shift that could slow American distribution even as Chinese competitors move aggressively through open-weight and lower-cost channels

To safeguard operations against future regulatory lockouts, enterprise technical leaders are moving toward model-agnostic fallback architectures.

By deploying proxy layers that can dynamically reroute critical production pipelines from proprietary APIs to locally hosted, open-weights alternatives, businesses can leverage top-tier capabilities without exposing themselves to single-point-of-failure vulnerabilities.

Fable 5 is officially back online, but the landscape governing its release has been fundamentally transformed.

🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI

1 July 2026 at 14:42

This episode has a fun personal twist: There’s a counterfactual world where I was employee #1 at Genesis Molecular AI,1 the company behind today’s episode. A certain introduction happened a few weeks too late and I had already happily signed at Atomwise2, another ML-for-drug-discovery startup. Same problem, different company. I was certain ML was going to transform small molecule drug discovery. Early results were underwhelming. Useful at times, but nowhere near revolutionary. In the last year I’ve seen signs that ML is finally ready to deliver on my convictions from a decade ago. Genesis is one of the places that might have finally cracked this problem. I was super excited to come full circle and catch up with co-founder Evan Feinberg and CTO Sergey Edunov.

If you are at all interested in small molecule drug discovery, we think you will find this fascinating!

In our nearly two hour chat we cover:

  • What is small molecule drug discovery, and why is it hard

  • Structure prediction as a hotbed of innovation in AI algorithms

  • How advances in AI elsewhere have enabled stepwise improvements in predictive power

  • How the community benchmarks are essentially calling AI slop good enough

  • The Genesis flagship model (PEARL) can routinely hit a threshold that is necessary for real-world applications

  • New agentic workflows enabled by these highly accurate models

Read on for more, and also some personal thoughts on the future at the end.

The coolest diffusion research is happening at Genesis

Sergey Edunov came to Genesis from Meta where he led Llama 2 training and Llama 3 pretraining. Sergey was a former physicist who thought he was done with physics after many years of training LLMs. Then, he discovered Genesis, and was blown away with all the novel architecture work they’ve been developing.

It probably surprises no one that modern LLM research has not resulted in fundamentally novel or exciting updates in architectures since almost the advent of the transformer — the entire field is using variants on the same idea that came out in the original “Attention is all you need” paper. Sure, some were quite useful (mixture-of-experts in particular allowed for the massive model paradigm we’re at today), but there was very little conceptually exciting.

“We sort of had to wait for the right primitive to get created, and that turned out to be diffusion… Actually, some of the most innovative diffusion research that’s happening in our field is happening in 3D structure prediction right now.” — Evan Feinberg

The field of 3D structure prediction on the other hand has been a hotbed of research. Genesis’ recent model PEARL (Place Every Atom at the Right Location) is able to understand protein flexibility, and model not just where the ligand goes, but also make small adjustments of the protein so that the two fit better than either alone. The field knew this was missing for a long time, but it was really hard to model until now.

Agentic Discovery

What makes this problem so hard? As Sergey points out, there are 10^60 possible drug-like small molecules. You’ll never be able to search them all, and trying to find the good ones is something like finding a needle in a haystack — except everything except your needle is dangerous.

“There are 10 to the 60 drug-like small molecules in the universe… it’s like finding a needle in a haystack, where everything except your needle is very, very dangerous.” — Sergey Edunov

“Or finding hay in a needle stack might be a more apt analogy.” — Evan Feinberg

Trying to solve the multi-parameter optimization problem is even worse. What makes a strong binder and a molecule with good “ADMET Properties”3 are oftentimes at tension with each other. For example, a good binder is likely greasy, but a greasy molecule is likely insoluble so it won’t enter the bloodstream and get to where it needs to go!

Genesis’ advances in generative AI have now pushed them beyond the threshold where they believe agentic drug discovery loops are finally possible. We all remember the early days of LLMs. They were great chatbots but terrible agents, as small errors compounded rapidly into uselessness. As LLMs got better, the usefulness of agents rapidly improved. Evan and Sergey argue that their models at Genesis recently passed a similar threshold. Their internal agentic drug-discovery system (code named SAPPHIRE) can now iterate like a chemist: look at and reason about poses, form hypotheses, read literature, use internal tools, create candidates for the next iteration. Combining this with automated lab partnerships like the one Genesis has with Incyte, we’re rapidly approaching a time of drug discovery agents running 24/7 making/testing new molecules. Exciting times!

Benchmark crisis: Everyone’s favorite benchmark is slop

One surprising point that isn’t talked enough about: the academic field of “co-folding” has settled on a benchmark value of “2 Angstrom RMSD” as a metric for a “good pose”. Evan does not mince words: this threshold is just bad. Perhaps even deceptively bad. For many strong binders, there’s a very clear pose, one that you can even directly resolve in the PDB electron density! And yet, with a 2Å RMSD threshold, you can get the pose quite wrong in ways that might even mislead a medicinal chemist. For example, flip around an aromatic ring, and everything looks reasonable, but you’re no longer modeling the right interactions.

Evan makes the strong claim that 1Å RMSD is really the threshold necessary to ensure the core of the molecule is sitting where it needs to be, and models all interactions.

“If your model is sitting at 1.8, 1.9 Angstrom RMSD, that’s slop, most likely.” — Evan Feinberg

As a simple example, he points out hydrogen bonds which are responsible for many of the most important interactions in protein-ligand systems. Hydrogen bonds only have a 0.6Å range to be valid! Clearly if you’re accurately resolving all H-bonds, you generally have to be doing much better than the 2Å threshold.

This is clearly a hard-fought lesson for Evan and Genesis. In their opinion, the community is stuck on these benchmarks because academics developing methods were not users. Evan does see signs of life, with the use of new metrics such as lDDT for co-folding. Hopefully soon the community can agree that “1.8Å RMSD is slop”, and start hill climbing on this much harder task.

For a more thorough exploration of the weaknesses in conventional benchmarks, see the PEARL technical report.

PEARL tops OpenBind

Which makes what happened next all the more striking. Near the end of the podcast, we talked about a recent “proof-is-in-the-pudding” moment for Genesis — evaluating their PEARL model on a recently released OpenBind benchmark. This benchmark featured 802 never before seen co-complexes on a target protein EV-A71. This target seems almost custom-chosen to give most classical docking methods a problem. When a ligand binds to the main binding site, the protein moves around to close off the path the ligand used to enter the binding pocket. This process, known as “induced fit” is notoriously hard for traditional methods to model. The tradeoff is easy to understand: treating the protein as a static structure, it becomes difficult to place a ligand in a binding pocket. Treat the protein as dynamic, and now you have to simulate complicated processes that take a long time to resolve.

PEARL was able to model the induced fit of the ligand without running long MD simulations. Across the different evaluation metrics, PEARL came out not just ahead, but oftentimes well ahead of any public model. A truly impressive result.

“Where PEARL was exceptionally good is figuring out how to move this loop. We are basically correct for every single pose.” — Sergey Edunov

Even more exciting, this was done without any fine-tuning, or using any data on the target or homologous targets — the template PDB was released after PEARL’s training cutoff.

Where does co-folding go now?

As someone who has followed or participated in ML techniques for protein-ligand interactions for almost a decade, I was genuinely impressed with the results that Genesis has released recently. This has been many years in development, and I’m sure Evan and the team had many sleepless nights trying to get to this point. I also think other teams are making similar progress — both Isomorphic and Deep Origin have released results that seem spiritually similar and combine computation, wetlab data, ML, to achieve genuine predictive power that seemed impossible a decade ago. Sadly, all of the above are closed source so there’s no way to honestly compare them. Looking at the results I think there might be a time in the not so distant future where we can consider protein-ligand binding “solved”.

I sincerely hope that the academic community can take inspiration from these developments. Once you know something can be done, it’s much easier to execute. Still, I believe that the key enabler in all of the above was the tight integration of ML, large-scale computation, and real-world drug discovery applications. Sadly academia is just not structured in a way that makes such a development easy.

With those parting thoughts, we hope you give the podcast a listen!

1

At the time called Genesis Therapeutics

2

Now called Numerion

3

ADMET stands for Absorption, Distribution, Metabolism, Excretion, and Toxicity. This set of about 30 properties all need to be optimized in order for a molecule to be considered a “good drug”.

💾

LLMs are stuck in a groupthink groove. This startup is trying to get them out.

Let’s start with a game. Open up your chatbot of choice—Claude, ChatGPT, Gemini—and type “Give me a random number between 1 and 10.” You’re going to get 7. Almost always. Now type “Another” and you’ll get 3 or 4. Type “Another” again and you’ll get 8 or 9.

That won’t work every time—but if it did, you may wonder if I have superpowers. I don’t.

The truth is that most large language models are stuck in a rut. They are far more predictable and far less creative in their responses than you might expect. That’s fine for tasks like coding or research, but groupthink is a problem when you’re brainstorming or planning your next vacation.

The Australian startup Springboards has a solution. It built an LLM called Flint, which has been trained to come up with a wider variety of responses than mainstream LLMs to open-ended questions such as “Where should I go in Europe?”

“Most language models are fighting hallucinations,” says Springboards cofounder and CEO Pip Bingemann. “We welcome them.”

Bingemann introduced me to the random number game when he first showed me his company’s new model. It felt like watching an illusionist with a deck of cards. “This is our sales trick, and it works every single time,” he says.

After ChatGPT and Claude both gave their 7s, Bingemann turned to Flint. It too came back with 7: “Aha, of course that was going to happen, but it’s okay—7 is a legitimate answer.” He restarted the session and prompted again: ChatGPT gave 7, Claude gave 7, Flint gave 3.7916.

Run your way

It’s not just numbers. When Bingemann asked ChatGPT and Claude to name a type of car, he predicted that it would be a Toyota or a Honda—and he was right. Flint came up with a Ford F-150. “There’s all this lost information that doesn’t get served up in these models,” he says. “They’re just as capable of saying a Buick or a Tesla. They just don’t—they’re biased.”

Bingemann sent one last prompt to each of the three models: “Give me a tagline for a campaign for New Balance running shoes. Just the tagline.” Claude: “Run your way.” ChatGPT: “Run your way.” Flint: “Built to last, run to win.” It won’t win any awards, but at least it’s different.

This weird limitation of LLMs is starting to get more attention. In November a team of researchers put out a paper, titled “Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond),” that exposed a remarkable degree of repetition not only in the answers from individual LLMs but between them as well. They found that different LLMs converged on very similar answers when prompted with open-ended questions.

It’s not clear exactly why this happens, but the researchers speculate it’s because most LLMs today are trained in similar ways on similar data to do similar tasks. The team won the best paper award at NeurIPS, a major AI conference.

When the researchers asked 25 different LLMs (including models from the top US firms as well as open-source models from China and elsewhere) 50 times each to write a metaphor about time, most of the 1,250 responses were a version of “Time is a river” or “Time is a weaver.”

(I asked some of my colleagues the same question and six people gave me six different answers. My highlight: “Time is a favorite sweatshirt, shaped by a lifetime of wear.”)

When you look for it, you see repetition everywhere, says Kieran Browne, cofounder and CTO at Springboards. “The way that most chat interfaces are designed, it makes it feel like you’re having a personal conversation,” he says. “I think most people don’t really realize the extent to which they are getting the same stuff as everybody else.”

Take another example: “What should I name my band?” Most models will say something involving “glass,” “neon,” “velvet,” or “static,” says Browne.  

When I tried it, ChatGPT spat out a list of 56 band names. At the top was “Glass Harbor.” Skimming through, I found “Static Empire,” “Neon Hearts,” and “Velvet Echo.” I asked Gemini; it gave me 15 suggestions, including “Static Horizon.”

Some of the suggestions looked pretty cool, though. ChatGPT’s “Sofa Astronauts” caught my eye, so I googled it—and found that a band called Sofa Astronauts already exists. 

(OpenAI says that training models to give reliable and coherent answers can lead them to converge around familiar, high-probability responses and that pushing harder for novelty can lead to weaker or less reliable responses. It also notes that the “Artificial Hivemind” paper studied models from 2024 that have since been updated.)

Creative catapult

Springboards has developed a tool backed by a selection of LLMs, including ChatGPT and Claude, that creative professionals in advertising or marketing can use to brainstorm ideas. The tool lets you drag around text produced by different models, picking the bits that you like and combining them into something new—in theory. Springboards is pitching Flint as an alternative model that users of its tool can select when looking for more variety.

Zoe Scaman, founder of the business strategy startup Bodacious and chief strategy officer at 77X, a direct-to-fan marketing platform set up by Luka Dončić of the LA Lakers, has been trying it out. “I find it really useful for throwing me in completely different directions,” she says. “I use it if I want to catapult myself all over the place.”

In one test, Scaman pitted Flint against Claude, Gemini, and ChatGPT by giving each of the models a classic MBA case study: How would you reinvent a finance company for today’s youth? The three mainstream models all went down the same path, she says: “You know, we need to teach financial literacy in a fun and funky way—well, that’s nothing new.”

But Flint came up with something different, suggesting that the whole concept of wealth accumulation should get a rebrand. “That was really interesting,” says Scaman.

She notes that Flint is still a prototype and doesn’t work all the time. “It sometimes falls over when you start pushing it too far,” she says. “But I think that the premise behind it is really powerful.”

Taking the temperature

Springboards built Flint on top of Qwen 3, an open-source model from the Chinese tech giant Alibaba. “We’re a small team,” says Browne. “Training a foundation model is not on the table for us. It’s just too expensive.”

Most LLMs have settings that let you adjust the level of randomness in their output. The most common is called temperature. “Obviously, that was one of the first things we explored, because that’s what people tell you: If you want more creativity, you turn up the temperature,” says Browne.

But changing those settings can also make models incoherent. Dialing up the temperature on one of OpenAI’s models to its maximum setting made it produce responses that switched from English into code halfway through a sentence, says Browne.

Springboards realized that parameters were blunt instruments for what it wanted to do. It does not make sense to dial up the randomness across the board; you only want to boost it at specific points in its output, he says.

For example, when you ask a chatbot “Where should I go in Europe?” the model only needs to tweak the randomness just before it names a destination, not for every word in its response.

To make Flint do this, Springboards trained its version of Qwen 3 to identify the points in its output where more variety was possible and fill those spots with words or phrases that were a little more random.

“Flint’s programmed to throw an oddball in. It’s more of an invitation to think wider,” says Maximilian Weigl, cofounder and chief strategy officer at Uncommon, a marketing firm. “That’s super interesting.”

Weigl’s team uses Flint alongside ChatGPT, Claude, and Gemini. “You can’t really create something boundary-breaking with tools that pull you back to the average,” he says. 

And yet Weigl notes that nine times out of 10 the average is fine. You don’t always need to reach for extremes with something like Flint, he says: “Most people are fine with good enough. They want to see mass-market familiar things.”

Weigl also cautions against using any LLM too much. “I have a big problem when people rely on the output from any AI, including Flint,” he says. “If I saw people on my team copy-pasting something from AI, I’d be like, ‘That’s not your job! Think, talk to other people, use your own voice.’”

For now, Flint is aimed at advertisers and marketers because those are Springboards’s customers. But Bingemann and Browne insist that a lack of variety is a problem for anyone using chatbots.

The idea is to give people the choice and leave it to them to decide if the result is good or not, says Bingemann. “Variety is great when you’re trying to spark ideas,” he says. “Let’s go down this route instead of letting the machines do it all and ending up in a gray, boring world.”

Warp CEO Zach Lloyd on why software factories are the next phase of coding

1 July 2026 at 14:28
Warp founder Zach Lloyd in the AI Engineer World’s Fair expo hall.

I’ve been covering Warp for a couple of years now, and its rapid evolution from a command-line interface tool to a software factory platform has been fascinating to watch. The company began in the pre-ChatGPT days, in mid-2021, as a Rust-based terminal. Then when AI hit, it turned into a terminal with integrated coding agents.

But the competition among CLI tools has dramatically increased in recent years, including from Claude Code, Codex CLI, and Gemini CLI — three products backed by massive tech companies. This likely led to Warp’s decision to open-source its core CLI tool in April this year.

I’m a Warp user myself, finding it a much more sophisticated tool than my native Mac CLI. But I also admire the company’s ability to adapt to the times — a trait I spotted in CEO Zach Lloyd during my first interview with him a couple of years ago. So I was keen to catch up with him at the AI Engineer World’s Fair this week, where he presented a keynote session on software factories, the new term for orchestrating a team (ahem, a factory) of agents.

Warp has a new agent orchestration platform called Oz. It’s the company’s answer to what Lloyd believes is an industry transition, from engineers working interactively with agents to automated systems that continuously triage, implement, review, verify and monitor software changes. Oz is intended to connect multiple models and coding harnesses across local environments and isolated cloud sandboxes, while fitting into tools developers already use.

I spoke to Lloyd just after he made his presentation on-stage, which you can view on YouTube — it’s a good primer to what software factories are. In our one-on-one discussion, we get into the reasons Warp made its software factory pivot, how Lloyd came up with the term (independently, it seems, from similar companies — like Factory), and why he expects most significant software projects to operate some form of automated factory within the next year.

From individual agents to an automated development loop

Latent Space: When did you first come across the term “software factory,” and what attracted you to the concept?

Zach Lloyd: I can’t remember exactly when I started conceiving of it in those terms, but it was within the last six months, as the ability to automate software development became more complete.

We started with more one-off automation: run an agent in the cloud. A lot of platforms began there. Then it became: run an agent in the cloud on a timer.

The next question was, what is the most valuable loop to automate? The answer is basically the main loop of software engineering: triage, specification, implementation, review, verification, shipping and monitoring.

We began building toward this cloud-automation vision about a year ago, before we started building Oz. Over the past few months, the industry has also begun coalescing around the ‘factory’ term. There is an entire software-factory track at this conference.

It is literally what we are gearing our product around. In the next version of Oz, you will set up your factory, see what it looks like and manage the factory floor.

But I don’t care that much whether the term sticks. The essential shift is from interactive development to automated development. “Factory” is a useful metaphor for that.

Building the factory around existing workflows

Latent Space: In your presentation, you showed a software-factory stack containing several of your own products. Is Warp’s plan to provide the tools that make up that stack?

Lloyd: Yes. When you enter Oz, our cloud-agent platform, you will be walked through setting up a factory.

You choose your repositories, the parts of the software lifecycle you want to automate, and the points where humans should be brought into the loop. Different organizations and codebases will have different preferences. Do you fully automate code review? Do you have humans review certain high-risk changes?

The system then starts creating the loop. It might pull issues from Jira or Linear, let people submit them through Slack or Teams, and allow developers to redirect an agent from GitHub.

What is interesting from a product perspective is that most of the factory is not necessarily a new interface. It is an integration into people’s existing workflows. That is how we are conceiving it, at least.

Why Warp is moving beyond the terminal

Latent Space: When I first wrote about Warp, it was building a modern terminal. Code is still important now, but increasingly it is being produced by agents. It looks like Warp has broadened its product vision accordingly...

Lloyd: One hundred percent. A good way to think about it is that the company’s mission has stayed the same since we founded it. It has always been about empowering developers and companies to ship better software more quickly.

The product has evolved tremendously. It began as a modern version of the terminal, before the current AI wave. The next iteration was a terminal with agents built into it, which we are still investing in and which we have now open-sourced.

But the world keeps changing. The underlying AI improves so quickly that my view of the future is what I described in the talk: the interactive component is going to become less important.

As a company, you will want a central place where software gets built and where you can measure the efficiency of that process. I’m not afraid to redirect what the product becomes. As the underlying technology gets better, companies that do not adapt are going to be left behind.

Factory engineering as a new discipline

Latent Space: The word “factory” may be off-putting to some developers, given its connotations with mechanism and rote work. What feedback have you received from AI engineers about this pivot?

Lloyd: The concept resonates strongly with the economic buyer — the person running the engineering team.

For an individual engineer, it can sound mechanized and uncreative. They may think: “I enjoy coding. Why would I want to work in a factory?”

One point I tried to communicate in the talk is that this will become a new engineering discipline. I think it can be extremely interesting if you view the job as meta-engineering: building the system that builds the product.

It uses many of the same problem-solving skills. You are asking why an agent performs one task well and another poorly. How should you adjust its feedback? What context does it need? How should the workflow change?

But, for better or worse, the power of these systems and their ability to accelerate software development are so great that writing everything by hand is not going to make sense for much longer.

Where forward-deployed engineers fit

Latent Space: Another trend at the conference is forward-deployed engineering, which often combines aspects of product management, consulting and traditional engineering. How does that fit into the software-factory model?

Lloyd: Standing up a software factory potentially involves integrating with a large number of existing systems, depending on the company.

The factory will work most effectively when it has context from those systems and is integrated throughout the organization’s workflow. A lot of forward-deployed engineering work in this area is effectively a transformation project.

It requires real engineering from someone who understands how to configure and deploy one of these systems. We do some of that, and some of our competitors do as well.

I don’t know what the final state will look like. Warp is approaching it more as a platform business than a services business. But there is certainly a business today in sending smart people into a company to transform its workflow using these products.

Warp as the test bed for Oz

Latent Space: I use Warp as my terminal, including for some coding tasks. What happens to the original Warp CLI product in the software-factory era?

Lloyd: When we open-sourced Warp, we put the repository under the control of Oz. We built a software factory around the open-source project, using our own factory platform.

We are still trying to improve Warp as much as possible. We are doing it with the community, and we are doing a lot of it with agents. In that sense, Warp is a test bed for the factory concept.

But it is also a product used by almost a million developers, many of whom rely on it as their primary development environment. We use it constantly ourselves, and we still have internal engineers whose job is to improve it. We are simply approaching that work with a factory mindset.

Gradual automation, not an overnight replacement

Latent Space: What do you expect the next year to look like, in terms of adoption of software factories?

Lloyd: This will not happen all at once. Engineers are not going to wake up one morning and discover that a software factory has replaced their jobs.

Companies will start with specific use cases, certain types of issues or lower-risk repositories. Those are places where they may be comfortable not having a human review every single line of code.

They will see how it performs. Then the engineering challenge becomes: instead of merging 20% of pull requests automatically, can we get to 30%, 40%, 50% or 60%?

There will still be a remaining percentage of work done by people because it is too difficult, ambiguous or dependent on greenfield thinking.

But I think this shift will happen over the next year. My prediction is that every significant software project will have some engine of code — something resembling a factory — continuously driving it forward.

It will become similar to GitHub or CI/CD: a standard part of how serious software projects operate. I would be surprised if that did not happen.

Start by automating the annoying parts

Latent Space: There are thousands of AI engineers at this conference. What should they do to prepare for this shift?

Lloyd: Instead of only building the product directly, try building some automation toward a factory and see what it feels like.

Suppose you want an agent to implement incoming user issues automatically. What is involved in making that work? What prevents you from adopting it?

Perhaps code review is the bottleneck. Perhaps the agent is making changes, but you cannot clearly see what it did. You only discover those problems by trying to build the loop.

Get out of the mindset of building everything by hand. Find an annoying part of your job and try to create a loop that handles it for you using a factory approach.

An agentic artificially intelligent X-ray scientist

Nature Machine Intelligence, Published online: 01 July 2026; doi:10.1038/s42256-026-01261-5

Chen et al. demonstrate an AI X-ray scientist that autonomously aligns single crystals at a real synchrotron beamline, showing how large language models can enable adaptive closed-loop experimentation at large-scale scientific facilities.

Privacy Act of 1974; System of Records

Pursuant to the Privacy Act of 1974, as amended, and Office of Management and Budget (OMB) Circular No. A-108, notice is hereby given that the Central Intelligence Agency ("CIA") is submitting to the Federal Register one (1) new System of Records Notice (SORN), CIA-46 Sexual Harassment/Assault Response and Prevention Office (SHARP) Records. This new SORN covers records related to CIA's assessment, processing, and tracking of sexual harassment and sexual assault inquiries, allegations, and reports, which includes case management, dispositions, guidance, and reporting to CIA leadership, stakeholders, and oversight bodies.
❌