Berner et al. show how to adapt popular neural networks into discretization-agnostic neural operators that learn from continuous scientific data, enabling scientific simulations that generalize more reliably across resolutions.
Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion.
IROS 2026: 27 September–1 October 2026, PITTSBURGH
Enjoy today’s videos!
NASA is considering a mission concept for an advanced, nuclear-powered rover to be deployed to the Moon’s South Pole as part of the agency’s Moon Base plans. The PROMISE (Polar Rover for Observation, Mapping, and In-Situ Exploration) mission concept relies on the Curiosity Mars rover mission’s testbed rover. Some elements of the Perseverance Mars testbed rover shown in this video could be used as well. As exact duplicates of Curiosity and Perseverance, the testbed rovers are equipped with flight-proven engineering systems capable of carrying technology as well as science instruments that would advance Moon Base efforts.
A Mars rover for the Moon? That’s some OPTIMISM right there.
The project explores soft, lightweight robots that can gently float around people in indoor environments and invite playful, affectionate, and everyday interactions. Unlike conventional drones, our robot is designed to be quiet, soft, touch-safe, and socially approachable. Through this work, we ask what future indoor companion robots might feel like if they were not rigid machines, but gentle floating beings that share space with us.
A couple of things from this new Figure video: Thing one is that the cart-pulling is a good illustration of how clumsy humanoid robots still are at basic tasks relative to humans. Thing two is that there are absolutely no humans anywhere near these robots. You can see one guy at 0:19, which I can only assume is an accident, because these robots are not safe to be around from an industrial safety perspective.
Welcome to Robot Park, where we’re building the future with Apollo 2. Robot Park is where Apollo learns today, getting the experience needed to make a difference tomorrow. Today we’re announcing Robot Park, our nearly 90,000-square-foot facility where Apollo 2 is collecting real-world training data needed to advance autonomous humanoid robots.
UBTech Robotics, the world’s first publicly traded humanoid robot-maker, has launched a humanlike robot that features lifelike silicone skin and “emotional AI,” as Chinese tech firms increasingly transition robots from the factory floor to the family living room.
Spherephones are redefining how we experience sound. Created at Georgia Tech, this wearable uses spatial audio to alert users to movement from every direction—including behind and below. Built for safer human-robot collaboration, the technology is expanding into gaming and accessibility applications. See how music is becoming a new language for awareness and interaction.
Humanoid robots are meant to carry out long-horizon autonomous missions in a world built for humans. This is hard. These missions consist of many steps, each of which requires them to perceive, navigate, and interact with the environment. This is exactly Flexion’s goal: building the general-purpose intelligence that turns any robot into a useful helper.
We’re introducing KinetIQ Ascend—our reinforcement-learning approach designed to reach 99.9 percent manipulation reliability at human speed and beyond.
Dr. Sebastian “Basti” Scherer has worked in field robotics since the first DARPA Grand Challenge in 2004. He runs the AirLab at Carnegie Mellon’s Robotics Institute and is the director of safe embodied AI at FieldAI. While much of the industry is focused on local skills like tabletop manipulation, Dr. Scherer sees the greatest value in solving dirty, dull, and dangerous tasks that require operating in uncertain environments where the robot needs to “just work.” When robots “just work,” they become less like robots and more like tools. “That’s the big challenge that we have to overcome,” he says. “And that’s the challenge that FieldAI is really primed to solve.”
Look, I really appreciate how valuable robots like ElliQ can be, and robots that do good work and offer a financial benefit are incredibly important, especially in the context of family care. But in my opinion, you really shouldn’t suggest that a robot with FaceTime or whatever is an equal replacement for in-person human companionship, nor should you suggest that AI can replace a human wellness coach. If you can’t afford those things, then sure, ElliQ can offer some of those capabilities in a very limited way, but that’s all.
Drawing inspiration from restaurant waiters in Morocco and Turkey, among other places, we equip a robot with a hanging tray to transport objects from one location to another without dropping them or spilling their contents. We incorporate this approach into an interactive robot waiter demonstration, which uses computer vision and visual servoing to steer toward a person with a raised hand to serve them.
It’s Los Alamos, so of course we have robots. Some work inside gloveboxes, while others probe unexploded ordnance in the field and aid with repetitive lifting, Doc Ock–style. Legend has it there’s a fro-yo robot in the cafeteria.
The night sky seems eternal and unchanging. But in cosmic time, nothing could be further from the truth.
Vasily Belokurov is one of three winners of the 2026 Kavli Prize in Astrophysics. The award is for uncovering fossil evidence of past galactic mergers that prove how the Milky Way evolved.
No matter the time or vantage point, from a pre-Neolithic cave to a post-lockdown London high-rise, the predictability of the night sky has always been humanity’s symbol of permanence and reassuring stability.
Yet this apparent calm is deceptive. Our galaxy, the Milky Way, emerged from chaos and turbulence, and its constellations are full of migrants, exiles and survivors. Right now, it has begun to stretch and distort again, pulled by a massive companion and heading for an inevitable collision.
How can I be so sure? As a galactic archaeologist, my job is to reconstruct the past of our galaxy and read the signs of its future.
Instead of digging through soil, I use the laws of dynamics and stellar evolution to sift through hundreds of millions of stars—searching for the most ancient and chemically peculiar among them, interpreting their orbits and piecing together the events that shaped the Milky Way. One ancient encounter left scars so deep that, billions of years later, they still define the galaxy around us.
I want to understand what governs the lives of these massive cosmic systems: which changes are nature—the slow internal evolution of a galaxy disk—and which are nurture, imposed by collisions and mergers.
Questions about the source of dark matter underpin it all. This is the invisible substance whose gravity holds galaxies together, but whose true identity remains one of the greatest unsolved puzzles in astrophysics.
The Milky Way is the one galaxy where stellar motions can be measured in extraordinary detail. This allows cosmologists including myself to construct our most precise map yet of dark matter: how far it reaches, how dense it is around the sun, what shape it has, and how smooth or lumpy it may be. If we can build this map in enough detail, we may begin to understand not just where dark matter is, but what it is.
A Cataclysmic Collision
Our work has been transformed by a revolution in open sky surveys. From 2000, the Sloan Digital Sky Survey showed what becomes possible when vast astronomical datasets are made public, enabling discoveries far beyond the goals for which the survey was first built.
And since 2014, Gaia, the European space telescope, has taken this transformation to another level by mapping the positions and motions of nearly 2 billion stars, turning the galaxy into a vast archaeological record. No ruins, no shards, and no bones—only stars that hold the clues.
The clearest giveaway that something cataclysmic took place long ago in our galaxy is the migrants we observe: stars that were not born in the Milky Way.
While native stars mostly travel together, circling the galactic center in the great rotating flow of the disk, migrants cut across that order. They slide past the locals, plunge into the inner galaxy, then fly back out to its outskirts, again and again.
These unusual orbits go hand-in-hand with unusual chemistry. Most of the migrant stars are less enriched in heavier elements than the locally born population. Their chemical composition is a sign of a slower rate of evolution that is typical of a dwarf galaxy.
This makes the migrants doubly valuable. They are both fossils of the Milky Way’s violent past and probes of its outer regions, traveling where the local stars rarely go.
How the Milky Way Was Rewired
One of the central ideas in the theory of cosmic structure formation is that galaxies grow hierarchically. Smaller galaxies fall into larger ones and are torn apart, leaving their stars behind as migrants.
In the Milky Way, the largest ancient structure of this kind is known as Gaia-Sausage-Enceladus. It is the remains of a vanished galaxy that collided with our own between 8 and 11 billion years ago (the “sausage” refers to a pattern in its stars’ motions).
The Milky Way also did not go through that crash unscathed. The collision rewired and reshaped it.
Some of these changes are easily visible in the data. Stars from the old disk were splashed into our galaxy’s halo, becoming exiles in the place where they were born. A new posse of star clusters were also acquired.
At the same time, we think something even more momentous was taking place. The encounter changed the orientation of the Milky Way’s disk, and its alignment with the dark matter halo.
Around the Milky Way, this dark matter forms a vast halo, much larger than the luminous part of our galaxy. We often imagine this halo as a sparse, round cloud, but Gaia has helped show this picture is too simple.
The dark halo can be stretched out of shape by a major encounter. Like a ship beginning to list, the Milky Way started to lean—not suddenly, not visibly, but over billions of years.
View of the Southern sky shows the Milky Way and (far right, close to horizon) two galactic neighbors, the Small and Large Magellanic Clouds. H.H. Heyer/ESO via Wikimedia Commons, CC BY-NC-ND
A New Galactic Dance
Unusually, compared with many galaxies of similar mass, the Milky Way was allowed ample time to recover from the shock of the “sausage merger.” No other cosmic cataclysm appears to have shaken our galaxy since, letting it settle into a quiet, uneventful life. That is, until now.
The Large Magellanic Cloud (LMC), currently our galaxy’s most massive companion, is already pulling at the Milky Way, disturbing its halo again. In an echo of what happened some 10 billion years ago, the Milky Way is being drawn into an accelerating dance with this neighboring dwarf galaxy, recoiling in response to the LMC’s approach.
This is a dance that only one galaxy is likely to survive intact. A new chapter of migration, survival and adaptation has begun.
None of this spoils the beauty of the night sky—it deepens it. The calm band of light above us is not a symbol of permanence, but the visible reminder of a long survival.
The Milky Way has been broken, rebuilt, and is now being disturbed again. Its stars remember the past; their motions reveal the future. What looks eternal is, in truth, a moment in a much longer story.
Most verticals aren’t clean, well-oiled SaaS databases; the reality is ugly documents, proprietary schemas, implicit workflows, and long‑running tasks that most general-purpose models struggle with.
This prompted construction project management company Trunk Tools to build a specialized, three-layer architecture — perception, semantics, agents — based on highly-detailed data to support high-accuracy, highly-relevant industry automation.
Their purpose-built stack has shrunk review cycles from months to days, prevented costly field errors, and given autonomous agents the ability to reason over millions of pages of documentation, the company says.
“We really set out to take the data from dispersed systems, pre-process it, structure it, go through our ontology into a knowledge graph, and then train AI models,” said Sarah Buchner, Trunk Tools' founder and CEO and a former carpenter.
For builders in other verticals, the company's approach could serve as a blueprint for transforming data chaos into agent‑ready, industry-specific workflows.
Where general-purpose LLMs break down on industry data
Foundation LLMs, while powerful, are optimized for breadth, not always depth.
“General-purpose LLMs are trained to be okay at everything, so they're weak at anything niche,” said Kriti Faujdar, a senior product manager working in AI infrastructure, agentic AI, security, and LLM platforms. For instance: Rare terms, domain-specific reasoning, the unspoken context that any practitioner “just knows.”
Web, app, and software developer Sébastien De Bollivier agreed that the biggest bottleneck is reliability on data that is “jargon-dense, abbreviation-heavy, and format-specific.”
“A GPT-4-class model can understand a French legal contract, but will fumble the specific article references practitioners need to cite,” he said.
Besides, the most valuable enterprise data never made it into pretraining anyway, Faujdar pointed out. It's sitting in internal systems and proprietary formats. “RAG helps a little,” she said. “But it's just giving better facts to a model that still can't reason properly in the domain.”
Pre-training on domain data is critical; enterprises should then fine-tune on good task examples and build their own evals. “A few thousand examples from real practitioners beats millions of scraped, noisy ones," Faujdar said.
Mixture-of-experts (MoE) can provide specialization without inference costs blowing up. Pairing RAG with fine-tuning also works well; RAG handles the factual long trail while fine-tuning fixes vocabulary and reasoning.
De Bollivier pointed to the advantage of hybrid stacks: A general-purpose model for reasoning and orchestration, a smaller fine-tuned model (or dense retrieval over a curated corpus) for domain-specific extraction. He advised: “Don't fine-tune to make the model 'smarter' about a domain, fine-tune to make it more reliable on the specific output format your workflow requires.”
The trades and construction are certainly industries seeing traction with these techniques, as are legal and healthcare, De Bollivier said. These verticals have “high stakes for errors plus standardized document formats, equaling clear domain-training ROI.”
One honest caveat worth mentioning, Faujdar said: Specialized models can often fall apart outside their domain, so they’re often not useful outside their expertise (unless they’re re-trained).
In highly-specialized domains like construction, “data dumps” into large language models (LLMs) don’t cut it, said Trunk Tools' CTO Amrish Kapoor. This is because most transformers are probabilistic models: When given an image, they report back that it is “probably” a tree, or “probably” a child playing next to a tree.
This makes them insufficient for high‑precision symbolic interpretation. For instance, in construction documents, a 2-millimeter-wide symbol has a vastly different meaning depending on where it’s placed.
Further, constrained by context limits, probabilistic models struggle with long‑term project memory. “I don't mean a context window of a few tokens,” Kapoor said. “I'm talking about long term memory that stretches across months and years, because this is how long some of these projects are.”
Instead, the company's three-layer system breaks workflows into:
Perception (reading and extracting data from messy docs like PDFs, drawings, or scans)
A semantic/graph layer (making sense of that data and understanding their relationships).
LLMs and agents on top.
Construction drawings are typically symbolic, Buchner said. A door isn't always labeled ‘door.’ Sometimes it's simply an arc on a wall that a trained eye learns to read based on years of practice.
“The perception layer is what teaches AI to read that language,” she said. The semantic layer then gives that information meaning; for instance, connecting the door to the drawing that details it, the spec that governs it, and the trade that installs it. This helps answer project engineers’ critical questions: Not "is there a door here?" but "does this door create a problem down the line?"
Particularly in construction, that shift matters because the cost of a problem compounds with time. “A conflict caught in design is relatively low cost to address,” Buchner said, “whereas the same problem caught in the field might cost tens of thousands of dollars.”
At a high level, the system identifies the document type and begins extracting information based on content (drawing, schedules, paragraph text). This data is then “transformed and augmented” in the platform, which triggers agentic workflows like knowledge graph relationships and end-user workflows.
For instance, an agent might review an architecture bulletin and produce a visual overlay comparing an older version and a newer version (flagging additions and removals), then generate written narratives that describe what those changes are in simple terms. This helps users understand what’s changed and coordinate with trade partners on updated pricing and change orders.
The scale of construction’s data problem
Construction workflows are “ripe with implicit assumptions and connections between data in its myriad of sources,” Buchner said. And the amount of unstructured data is “humanly impossible” to process or make sense of.
Buchner estimated the average high-rise building generates about 3.6 million pages of corresponding documentation. “If you print it into a stack of papers it would be as high as the building itself.”
All three layers of Trunk Tools' stack — perception, semantic, LLM — are trained on “very specific datasets” from customers with “explicit permissions” and auto‑labeling/IP, Kapoor explained. Customers who don’t want Trunk training on their data can opt out.
Data is deidentified and aggregated, and Trunk Tools also collects “tons more” labeled data through other pipelines like 3D building information modeling (BIM).
The company says it only ships agents that achieve around 95% accuracy. The team maintains continuous evaluation pipelines based on ground truth data from customers and experts. They also employ an LLMs-as-a-judge model.
“This notion of an LLM as a judge is to score how well you're doing, both subjectively as well as objectively,” Kapoor said. Objectivity can be an easy ‘right’ or ‘not right,’ but subjectivity requires more nuance.
For instance, when creating an email or narrative or explanation, an LLM as a judge framework can create a composite score, or a numerical value that aggregates different metrics and tests a model's performance or risk.
There can be challenges, though, particularly with latency, Buchner noted; any time the reasoning capacity of underlying models increases, the risk of latency goes up, too. Trunk Tools maintains a set of evaluation criteria to objectively measure latency whenever changes are made to underlying infrastructure, agents, and API calls.
Then, “before we release to customers, we ensure marginal changes to the end-user experience are well worth the performance enhancements,” Buchner said.
From 60 days to 10: the measurable payoff
Trunk Tools' platform powers seven AI agents purpose-built for construction, such as analyzing request for information (RFI) responses, overviewing bids, or reviewing drawings and submittals.
The submittal agent, for instance, flags missing, conflicting, or noncompliant information in product specs and RFIs. While it’s an essential step in the construction process, “it's a super annoying workflow,” Buchner said, because human reviewers have to compare documents “with a bunch of other parts of documents.”
But the agent is able to do this in seconds, and Trunk Tools says it has reduced submittal cycles from 50 to 60 days to 10, “which has massive schedule and financial implications.”
The company is now at a place where these agents are communicating directly with each other, which is “quite exciting,” Buchner said. So, for example, one agent will review an architectural drawing for accuracy, then autonomously hand it over to agents handling RFIs and asking follow-up questions.
“If the drawings have problems, the RFI agent is taking over and is actively reaching out for clarification,” Buchner explained.
Trunk Tools says its customers report savings of 20 to 40 minutes per field question. Buchner said that users in the field know better than anyone how much of a “time suck” it is to go back and forth from office trailers, dig through project documents in scattered systems or printed PDFs, reconcile discrepancies, and return to coordinate with trade partners.
The company says its customers report these additional outcomes:
Average 8 minute time savings for single-document retrieval (status checks, location lookups, quantity queries).
Average 20 minute time savings for standard referencing (cross-referencing 2 to 3 spec sections to form an answer.
Average 40 minute time savings for multi-document research (listing and filtering queries, mapping relationships, analyzing RFIs and submittals across 4 to 6 documents).
Average 75 minute time savings for complex tasks (creating RFIs and other communication materials, deep cross-referencing across documents, change tracking).
In one instance, the company's drawing review agent flagged that a structural beam had been moved up 8.5 inches. However, this was not documented by the architect. If the change hadn’t been caught, the project manager would likely have had to strip out and reinstall the right size beam, Buchner said. This rework would have added $10,000 or more to the budget, and “certainly there would have been implications on the schedule.”
Buchner also pointed to other examples: an agent flagged $60,000 in exaggerated pricing with no justification from landscaping subcontractors; identified a fireplace that needed to be sealed prior to drywall installation, saving around $100,000 in labor, materials, and delays; and called out that an electric door required a panel that wasn’t included in electrical drawings.
Learnings for other industries
Trunk Tools' approach to building agents is applicable to any vertical working with high volumes of unstructured, industry-specific data.
Builders working in specific verticals must understand the industry’s specific data challenges their end users face and build technical infrastructure that can transform unstructured data into something an “LLM can traverse and understand,” Buchner said.
“Only then can you build the connections between data points that ultimately feed agentic workflows.”
A lot of money is being invested in foundational models, so enterprises should build modular systems that can leverage the strengths of various models as they continue to improve, Buchner advised.
Then, “build your technical advantage where the generic models are not investing and not performing well,” she said.
This episode of This Week in AI arrived at a moment when the AI infrastructure most teams take for granted suddenly looked a lot less stable. Andreas Welsch, founder and chief human AI officer at Intelligence Briefing, was joined by Matt Palmer, head of developer experience at Conductor and developer educator on LinkedIn Learning, to […]
One of the highlights of the final day of the AI Engineer World’s Fair was a debate about loops. It nicely captured an argument running through the whole conference: are autonomous software factories viable now, or is the engineering discipline lagging behind the ambition?
Allie Howe from Keycard was the moderator and she opened by asking, “is there or is there not a delta between the hype behind loops and what actually works in practice?”
The pro-loop case was presented by Geoffrey Huntley, creator of the Ralph Loop, and Keycard CEO Ian Livingstone. Huntley opened by saying loops are already here. “It’s inevitable, it’s here to stay,” adding that “I don’t see myself going back to writing code by hand.”
Livingstone said that verifiability is ultimately what it’s about — and you can achieve that with any code, regardless of how it was produced. He also pointed out that loops have always been a core aspect of software development:
“A loop is at the core of ‘I try something, I learn something, I apply something.’ And all we’re really talking about is how quickly we can expedite that process.”
On the skeptical side were Dex Horthy from HumanLayer and Greg Pstrucha from Subroutine. Horthy began by noting that he wasn’t anti-loops. “The basic take here is not whether loops are good or bad,” he said, noting that “Kubernetes is actually built on loops — built on control loops. But they’re deterministic loops.” Horthy’s issue is that “the hype is outrunning the discipline.”
“I haven’t seen proof that we are at a point where we can just step up an abstraction level,” Horthy said, referring to agents controlling the coding. “I actually think we need to step down an abstraction level, if anything.”
Pstrucha was mainly concerned about the economic viability of agentic loops, which he said wasn’t sustainable. You can’t “orchestrate your problems away by buying more tokens,” he said.
“[We’re] kind of like locomotive engineers now. That’s our job: to keep the locomotive on the rails.” - Geoffrey Huntley, loops advocate
Huntley then offered this wonderful analogy for loopmaxxing: “[We’re] kind of like locomotive engineers now. That’s our job: to keep the locomotive on the rails.”
The discussion turned to software factories, the metaphor that has really taken hold of the industry. Horthy worries that when everything is automated in a factory-like agent environment, “you never touch the problem.” So instead, he advises to start small and iterate with agent loops — to “build up intuition” and not try to automate end to end from the start.
Even Huntley recognized some of the dangers in loops. He said that software factories represent where we are headed in the future, but cautioned that it’s not yet solved in the market. “This is frontier thinking,” he said.
At the end of the hour-long debate, Howe polled the audience to ask which side ‘won’. Ironically, this resulted in a human failure: the stage lights were too bright for Howe or any of the debate participants to see how many hands were raised. If only an agent was in charge of dimming the lights.
Anthropic’s next big thing: Claude Tag
Perhaps one example of a company moving to a software factory model is Anthropic. Mike Krieger, one of the co-founders of Instagram back in Web 2.0 and now Head of Labs at Anthropic, was interviewed by swyx in one of the morning sessions.
Krieger talked about Claude Tag, Anthropic’s internal model which the company announced to the world last week. He described Tag as more delegated, asynchronous and proactive than Claude. It perhaps suggests what an early software factory looks like in practice — not agents replacing a team, but multiple people delegating responsibilities to a system like Claude Tag.
Mike Krieger talking with swyx at AIEWF today.
“Most usage is actually much more delegated,” he said regarding his team’s usage of Tag. He gave an example of how they instruct the agents: “Don’t just fix this bug. Now you are responsible for this part of the codebase, and I want you to monitor this feedback channel and proactively take on tasks.”
“That’s really changed how we operate currently,” he continued. “It’s much more this multiplayer, async, proactive way.”
However, he also indicated there are some negative consequences to becoming more automated. He noted that his team is “bottlenecked on reviews” and on the “human ability to fully conceptualize what we’re doing.”
2026 AI Engineer Survey
Back to the current reality for most AI engineers. This morning, Barr Yaron from Amplify presented her annual survey of the industry.
According to Amplify’s data, 95% of respondents now use agents — roughly double last year’s share. Among teams using agents, 89% said those agents could write data, up from 52% the previous year.
“Agents are no longer reading, summarizing, drafting,” Yaron said. “They’re taking actions inside the systems.”
Barr Yaron presenting her AI engineering survey.
The controls, however, remain comparatively primitive. Human approvals and permissions were the two leading safeguards, followed by a scattered collection of task decomposition, retrieval, memory and sandboxing techniques.
“Nobody has settled the control layer for agents,” Yaron said.
Cost is also a concern. Forty percent of respondents said that AI costs regularly limit how ambitiously they use AI, while another 36% said it sometimes does. Token usage is now the second-most monitored production metric, behind quality.
The survey captured the conference’s central contradiction. AI has made experimentation cheaper and enabled teams to produce more software, but 59% of respondents to the Amplify survey fear that today’s AI-generated code is creating long-term liabilities.
Closing keynotes
The final sessions of the conference appropriately took us back to thinking optimistically about AI technology — about building with it. After all, that’s why the AI Engineer World’s Fair exists, and it’s where the fun is!
Theo Browne showcased several software projects he had built, or was still building, with AI. His point was that the scale of what an individual developer can realistically attempt has shifted. “What used to be a startup is now a side project,” he said, while projects he would once have dismissed as “too big” are moving within reach.
Garry Tan, president and CEO of Y Combinator, followed by giving that optimism an organizational form. The fastest-growing founders YC sees, he said, are “not treating AI as autocomplete, they’re treating it as a workforce.”
Garry Tan at AIEWF.
Tan’s closing prescription was: “Build an AI-native company, not a company that just uses AI.”
The debates during the week showed how much engineering remains before the AI-native vision is viable for all. But the closing keynotes offered a reminder of why the engineers who attended this conference are pursuing it: they just want to ride those locomotives!
Two-thirds of enterprises have hedged their AI model strategy, and the past few weeks of controversy around Anthropic’s Claude Fable 5 model showed why that posture has gone mainstream.
On June 12, a U.S. export-control order pulled Anthropic's Claude Fable 5 — the most capable model on the market — offline for every customer, with no warning and no timeline. It returned this week wrapped in tighter safeguards, after China's Z.ai released its open-weights GLM-5.2 into the vacuum. New VentureBeat Pulse Research, which surveyed 145 enterprises across these last few weeks, shows that two-thirds had already hedged their model strategy before the order came down: 51% blend closed frontier models with open-weight models deployed on their own infrastructure, and another 16% are moving core workflows off closed APIs entirely. The remaining third was all-in on closed ecosystems when the lights went out.
The blackout put a spotlight on vendor dependency, by showing what happens when the model you rely on disappears. But vendor dependency is only the most visible piece of a deeper problem: Most enterprises lack the monitoring to know when an AI system they've put into production stops working correctly.
Just 1 in 10 enterprises has automated monitoring that would catch an AI model drifting, misbehaving, or failing in production. Roughly a quarter would learn of a production failure only when end users — internal or external — report it, or lack the visibility to detect it at all. And 79% of enterprise organizations have already taken a real financial or operational hit from autonomous agents — most often shadow AI, unauthorized agentic work run by enterprises' own employees on corporate credit cards, outside anyone's oversight.
We call this the “Control Gap,” or the distance between how aggressively enterprises are deploying AI and how little of it they can see, own, or govern. June’s blackout turned this into a live stress test.
About this data: VentureBeat Pulse Research surveyed 145 qualified respondents at organizations with 100 or more employees in June 2026, with fielding spanning the Fable 5 blackout that began June 12. The sample is self-selected and directional: 41% work in technology/software, 20% are consultants or advisors, and the respondent base skews senior and technical — CIO/CTO/CISOs (18%), directors of engineering/IT (14%), enterprise architects (12%). More than half of the respondents were from companies with 2,500 employees or more.
While our sample is not huge, what you can trust more than the exact percentages is the pattern: Every question in the survey, independently, points the same way, with deployment running ahead of governance, visibility, and cost control.
How the Fable 5 export order rewrote enterprise AI risk
Fable 5 launched June 9 to immediate acclaim — and sticker shock, at $10 per million input tokens and $50 per million output. Three days later, the U.S. government issued an emergency export-control directive barring access by foreign nationals. Anthropic, with no way to verify nationality in real time, suspended the model for everyone.
June added the harder lesson: The model your workflows depend on can vanish overnight, by government order, through no decision of yours or your vendor's. And Chinese companies like DeepSeek were releasing hugely disruptive, powerful models, driving down costs to a fraction of Western ones.
Brian Craig, senior director of architecture at Liberty IT, the Ireland-based engineering arm of Liberty Mutual, one of the world’s largest insurance companies, saw both lessons collide in real time. Craig is Irish, which meant the export order hit him directly as a foreign-national user.
Onstage at VentureBeat's AI Impact event in New York on June 24, mid-blackout, I asked him about it. "Fable arrived, and immediately you saw the sticker price of using it, and you went, 'Ooh, goodness, it better be really good,'" Craig said. "But luckily enough, we didn’t get to use it enough to get to fall in love with it." Then it was gone.
The hedge was already built before the blackout hit
Craig's company was built to route around exactly this kind of disruption. Liberty IT runs what it calls an AI backbone — roughly 50 components spanning security, governance, observability, and orchestration, each independently replaceable.
"You can't lock in right now in one vendor and even one framework," Craig told the room. "You need to keep being able to have the flexibility with that backbone to be able to hook into different models, different vendors, depending not so much on who's the flavor of the day, but on what you can feel confident about for the next six months."
The survey shows Craig has plenty of company. A 51% majority of enterprises run a hybrid posture — closed frontier models for general reasoning, open-weight models deployed locally for specialized execution — and 16% are making a hard pivot, moving core workflows onto open weights running on their own hybrid or private cloud. The 32% holding a closed commitment are candid about why: The operational overhead of self-hosting still outweighs the savings for them. After June, that calculus has a new variable in it.
Defection is now the active posture, and the target may surprise you. Asked which primary AI vendor they are most likely to downsize or phase out over the next 12 months, respondents named Microsoft first at 30% — most citing cutbacks to Copilot and Azure AI frameworks in favor of direct model access — ahead of the 28% who plan to trim no vendor at all. OpenAI drew 21%, largely on pricing volatility, with Anthropic at 15% and Google at 6%. No vendor faces an exodus. But loyalty by inertia has ended: Among these enterprises, actively cutting at least one provider is now more common than expanding across all of them.
Just 1 in 10 enterprises would catch a failing production model automatically
How would an enterprise know if one of its production AI models was drifting, behaving unsafely, or failing to complete tasks? We asked directly. Forty percent say they are very confident they would detect it. The question also asked what that confidence rests on, and respondents split into two camps: 30% rely on humans reviewing critical AI outputs, and just 10% — 14 of the 145 organizations — have automated monitoring and alerting running against production systems. The remaining respondents hold weaker positions still: 32% expect to catch most issues "eventually," 19% say they would likely hear about a failure from end users first, and 8% report no systematic visibility into production AI behavior at all.
That distinction matters because the two approaches are very different. Human review may seem like the gold standard, but it only reaches the outputs someone designates as important for such a review — and it happens at the pace humans can move at, with the inconsistency any manual process carries. Automated monitoring watches everything the system produces, continuously, and flags anomalies as they happen — for the same reason enterprises stopped depending on manual checks for uptime and security a decade ago.
As agentic workloads multiply output volumes far beyond what any review team can read, the manual approach starts to fall behind. The leaders at our June 24 event in New York treat human review as a designed control with automation underneath it. "Nothing gets deployed into production unless it's a human actually reviewing it and signing off," Craig said of Liberty's agentic software factory, where planning, coding, testing, critic, and librarian agents ship features from epic to production.
"It always has to be risk-based. That's why we work for an insurance company." Todd Johnson, the Morgan Stanley managing director who runs agentic AI across the bank's end-of-day P&L controller process, described the same principle from finance: "One of our strong principles in our AI governance generally is that there always has to be human accountability, even if there's a degree of automation." VentureBeat covered Morgan Stanley's new results around its P&L resolution agent system separately.
Liberty Mutual and Morgan Stanley chose manual sign-off deliberately, layered on top of observability, identity, and governance infrastructure. Whether the human-review camp has similar infrastructure underneath is more than a single-select question can establish. The 16% who separately named missing observability tooling as their biggest governance barrier are the ones saying outright that it hasn't been built.
The top governance barrier is organizational: no single owner for AI across platforms
Why does the AI visibility tooling never get built? The respondents' answers suggest it is an organizational shortcoming. The single most-cited barrier to governing AI across platforms is the absence of a single owner or accountable team, at 32%. Vendor opacity follows at 25%, missing tooling at 16% — and a lack of talent lands dead last at 5%.
The skills exist, but the organizational mandate does not: Only 38% say a central team actually governs AI behavior across their platforms today, 21% say ownership is unclear or actively contested between teams, and 17% say no role holds formal accountability at all.
The AI surface being governed makes the vacuum worse. Fully 85% of enterprises run two or more platforms each claiming to be the "primary" AI layer — ERP, ITSM, productivity suite, data platform, each with its own AI, its own controls, and its own assumptions. 36% describe an open contest between four or more. Just 8% have consolidated to one. Asked in a free-text question what one thing they would fix, respondents converged from different directions on the same answer: a single accountable owner, and a control plane that abstracts cost, drift, and model choice away from the end user.
79% have already paid for an agent control failure — led by shadow AI
The cost of the vacuum is showing up on corporate cards.
Asked to name the most severe financial or operational control failure they have experienced from autonomous agents, 49% of enterprises cite shadow AI — departmental teams running unauthorized agentic pipelines on corporate credit cards, bypassing central financial oversight entirely. Another 25% have been hit by an infinite-loop bill, an uncaught recursive workflow racking up thousands in token costs in a single incident, and 6% by an agent that degraded production databases with unthrottled queries. Only 21% report guarded stability, with hard token throttling and budget caps at the infrastructure layer. Add it up: 79% of these enterprises have already paid for an agent control failure in real money or real downtime.
Finally, the economics of tokens suggest the pressure will keep rising. Per-token inference costs are falling 70 to 80% a year, and agentic workloads consume 100 to 500 times the tokens of the LLM tools they replaced.
Brian Gracely, senior director of portfolio strategy at Red Hat, told our New York audience the answer starts with right-sizing: "If I'm simply trying to resolve an insurance claim, I don't need to know about the history of Western civilization in my model. I don't need to know soccer scores."
Enterprises are pairing smaller, specialized models with semantic routing, he said, so the platform decides which requests genuinely need frontier-scale reasoning — and which are burning premium tokens on commodity work. (One adjacent data point from the survey underlines the appetite for pragmatism: 73% of enterprises report little or nothing to show for their custom fine-tuning investments of the past 18 months — a reckoning we'll examine in its own report.)
The bottom line: Replaceability is spreading faster than ownership
The survey describes enterprises moving fast on AI with weak controls underneath. 58% are adding more AI initiatives than they retire. 85% run multiple platforms that each claim to be the primary AI layer. Three times as many enterprises rely on human review to catch a failing production model as have automated monitoring in place. And 79% have already paid for an agent control failure — most often unauthorized agent spending on corporate cards, outside IT's oversight.
On one problem, enterprises have clearly adapted: model dependency. Two-thirds hedge their model strategy, either running open-weight models alongside closed ones (51%) or moving core workflows off closed APIs entirely (16%). The Fable 5 shutdown showed the value of that position — the hedged companies could route around a model that a government order made unavailable overnight.
The remaining problems are internal, and no purchase fixes them: 32% name the lack of a single accountable owner as their top governance barrier, and 17% say no role holds formal accountability for AI at all. Assigning an owner costs nothing and requires no vendor. It still hasn't happened at most of these companies.
Our coming Q3 wave of research will measure whether June changed this — whether enterprises assigned owners and installed automated monitoring, or just added a second model and moved on.
The themes in this report — agent orchestration, governance, and cost control — are the agenda at VB Transform, VentureBeat's flagship event, July 14-15 at Hotel Nia in Menlo Park, with technical leaders from Visa, GM, Waymo, Intuit, Instacart, LangChain and others. Details and registration here.
Disclosure: VentureBeat's June 24 AI Impact event in New York was sponsored by Red Hat and Intel. Sponsors have no input into VentureBeat Pulse Research survey design, findings, or editorial coverage.
Andrew Qu is Chief of Software at Vercel, where he works with the CTO across internal engineering, product experimentation and emerging technologies. He has built libraries for MCP, created skills.sh and led the development of eve, Vercel’s framework for building agents.
In this interview with Latent Space, Qu explains why agents represent a new form of software, what Vercel learned from building its own, and why Vercel itself is turning into an agent!
From web applications to agents
Latent Space: What does a Chief of Software do at Vercel?
Andrew Qu: My role is pretty unique. I work with the CTO to ship impact in any way, shape or form. It’s a mix of internal engineering, external experimentation and staying on the frontier by building things.
That means building new libraries and frameworks and showing people how to do things for the first time. I built an MCP library that made it easier to create some of the first MCP servers, and I also built skills.sh to make agent skills easier to discover and use.
Latent Space: How did Vercel evolve from focusing on web development to investing heavily in agents?
Qu: Vercel’s origins were about making it easy for developers to ship websites and web applications. More recently, we’ve seen a shift from people building pages to people building agents.
While building our own agent in v0, our vibe-coding product, we ran into a lot of paper cuts that existing tooling did not solve: switching models or providers, adding fallbacks and making runs resumable.
We turned those solutions into reusable libraries that could support v0 and also help customers build their own agents. Over time, we accumulated a set of primitives and decided to assemble them more cohesively. That became eve.
Why eve became necessary
Latent Space: How did you reach the point where Vercel needed a dedicated agent framework?
Qu: About a year ago, I started working toward putting an agent on every desk inside Vercel. That led me to build a successful data agent, and along the way a number of best practices emerged: filesystem agents, skills, compaction and subagents.
These were all things I wished had come out of the box. Eventually, we asked: what if there were a prescriptive way to do this, so other developers did not have to go through the same exploration? That is where eve came from.
Latent Space: Are agents simply another kind of application, or a genuinely new form of software?
Qu: I think agents are a new type of software. They are not as predictable as web applications. The infrastructure can look similar, but the interaction, interface and outputs are much more dynamic.
That changes how you build them. You need different primitives for context, tools, resumability and long-running work.
Latent Space: What kinds of problems are particularly well suited to agents?
Qu: We see a lot of business agents. Internally at Vercel, we use them for repetitive work ranging from a first pass at legal contract redlining, to marketing retrospectives and identifying people to contact, to writing queries against our data stores.
A good candidate is often a repetitive task that still requires some reasoning. It is not just fixed automation, because the system has to interpret the situation and decide what to do.
Building effective agents
Latent Space: When should an agent work autonomously, and when should a human remain in the loop?
Qu: I don’t think the future is all autonomous loops, and I don’t think it is all human-in-the-loop. It is about choosing a feedback cycle that fits the task.
If the task is well defined and you know what the final output should look like, it can be reasonable to let a loop continue until it is done. For more careful or surgical engineering work, you should check back in and make sure you are steering the model correctly.
Latent Space: Your approach evolved through prompting, bespoke tools, coding-agent harnesses, filesystem agents and skills. What was the main lesson?
Qu: We are still figuring out what makes an agent productive. Along the way, we have been collecting these primitives and bringing them together in eve.
There will be more to add as best practices emerge. A year ago, we did not know sandboxes would become so important, or how much demand there would be for secure code execution and long-running jobs. As we learn more from production, there will be much more to build.
Latent Space: Is Vercel creating an end-to-end agent platform comparable to the one it built for web development?
Qu: Yes and no. We value partners that provide specialized parts of the agent lifecycle, but we also want it to be very easy for developers to get started.
If you deploy eve to Vercel, you get observability and evaluations out of the box. We want to make that experience more comprehensive while making it easy to integrate with partners rather than owning every component.
Skills and current knowledge
Latent Space: Why have skills become so important?
Qu: Skills are useful as portable, on-demand knowledge. Models often contain outdated information. For example, they still sometimes recommend Vercel Postgres, even though we deprecated it years ago in favor of our marketplace.
A skill can tell the agent that Vercel Postgres is deprecated and steer it toward the current approach. Until companies can audit and update every old piece of content, skills provide a way to forward-correct the model.
I would recommend publishing skills for the latest version of your product. But companies should also audit their existing content, identify what is outdated and update it or add clear notes.
An agent-readable web
Latent Space: How will websites evolve as more traffic comes from agents?
Qu: We have published reports showing bot traffic rising while human traffic is stagnant or declining, even as impressions increase, because agents and bots are hitting websites more frequently.
The future of the web is therefore to be as accessible to bots and agents as possible, so they can learn about your product and use it successfully.
At Vercel, we already detect when an agent makes a request and serve Markdown directly. Instead of forcing it to process HTML designed for a visual browser, we provide a format that is easier to read.
Latent Space: Does that mean one experience for humans and another for agents?
Qu: I think so. Humans may continue to receive the visual site, while agents receive a more structured, machine-readable representation. We are already doing that today.
What comes next
Latent Space: What problems are you most interested in solving next?
Qu: One of the things at the top of my agenda is multiplayer agent development. Whenever a team collaborates, people struggle to share context.
I may have techniques for getting a front-end interface right on the first attempt, but another person may not know them. I am interested in how we can share that context between teammates and allow them to contribute to it.
Latent Space: Will agents become a separate application category, or a standard capability built into most software?
Qu: It depends on who you are and what you are building. For Vercel, Vercel itself is becoming an agent. We have an agent on the website, in Slack and in the dashboard that can do things on your behalf.
Other companies will ship agents as standalone products. For us, agents are tightly coupled to everything we build. We want the entire platform to be agent-friendly — and, in many ways, to make the platform itself an agent.
AI has transformed how organizations operate, driving unprecedented levels of productivity and innovation. However, AI adoption can be impeded by concerns...
AI has transformed how organizations operate, driving unprecedented levels of productivity and innovation. However, AI adoption can be impeded by concerns surrounding data privacy, sovereignty and how to secure data while it is in use, or during inference and engagement with AI models. NVIDIA Confidential Computing (CC) was engineered to be a secure and performant solution for the era of agentic…
Adobe Principal Scientist Carlos Sanchez at AIEWF.
For as long as I can remember (and I managed websites in the dot-com period), “personalization” has been a holy grail for websites. But up till now, that’s typically meant selecting from a predefined set of options. A retailer might recommend an item based on a previous purchase, or place a visitor into one of several audience segments — that’s been the extent of personalization.
Adobe Principal Scientist Carlos Sanchez is exploring a more radical possibility: what if the website itself could be assembled around the needs of each visitor?
At the AI Engineer World’s Fair in San Francisco, Sanchez demonstrated what Adobe calls an “agentic site” — a web experience that interprets a visitor’s intent, retrieves relevant material from the company’s existing content, and composes a personalized page in real time.
Adobe calls this approach an “audience of one.” Sanchez’s larger point was that the technology is no longer hypothetical.
“Many people don’t even think it’s possible to generate a web page on the fly,” he told Latent Space after his session. “People think it is future-looking. No, you can do this. It’s not the future, it’s the present now.”
From personalized components to personalized pages
During his presentation, Sanchez demonstrated a site that used the visitor’s browsing behavior and search queries as signals. The system grouped those signals into an intent category — such as exploring, researching or preparing to purchase — and then used an LLM to assemble a page suited to that intent.
In one example, a visitor interested in camping received a version of a coffee-machine site whose copy, product selection and supporting content had been reorganized around making coffee outdoors.
Sanchez also showed a more open-ended interface in which someone could enter a query such as “Europe AI conferences” and receive a page composed specifically around that request.
“We call this ‘audience of one,’ because the idea is to personalize the site in real time based on the user accessing it and what the user is doing,” Sanchez said.
The idea is that the site’s existing content is the grounding corpus. Adobe’s system retrieves from that material rather than asking an LLM model to invent an entire experience from scratch.
For AI engineers, one potential constraint is latency. In his session, Sanchez said that Adobe evaluates models not only for accuracy, but also for speed: “We don’t want the site generation to take more than one or two seconds.”
Sanchez says the economics are already becoming plausible. He estimated the current inference cost at “one to two cents per page.”
“But our point is also this is only going to get cheaper,” he said. “This is where we are today. In six months, who knows where we’re going to be.”
AI makes it easier to build, but harder to choose
Adobe has not yet broadly deployed these experiences on production customer sites. Sanchez said the company is presenting the concept to customers and looking for organizations willing to experiment.
Commerce is an obvious initial use case, because personalization can be connected directly to conversion. But the opportunity is not necessarily limited to retail. “It could work for other things — anything that needs more conversion and has a big matrix of user types or personas,” he told me.
Still, Sanchez acknowledged that he’s unsure if agentic sites will become a widespread reality.
“With AI, it’s very easy to build things, but it’s hard to know what to build,” he said. “We build things and then we find the customers.”
It’s not just Adobe feeling the uncertainty around its ‘audience of one’ concept. Website owners are currently evaluating all kinds of AI functionality: chat interfaces, structured content (like WebMCP), generative UI, personal agents, and more. Not to mention trying to find ways to bring users in from third-party AI platforms.
“I think it’s a combination of all these crazy different ways,” Sanchez said. “You are in a chat, I want to show UI, I want to get you to buy something. Then you’re in a site, I want to steer you this other way. Maybe you’re in an OpenAI chat and I want to bring you into my site. Everybody’s trying to figure this out on the marketing side.”
A web built for humans — and agents
Of course, websites in 2026 and beyond won’t just be personalized for human visitors.
As personal agents become more capable, a user may delegate some purchases or research tasks entirely. The agent could arrive carrying a much richer expression of the user’s preferences than the destination site could infer from cookies or recent browsing behavior.
Sanchez expects websites to evolve for both kinds of visitor. “Whether it’s going to be two versions [of a website] or not, that may be blurry,” he said. “But obviously, you’re going to have to target both.”
Also, not every transaction will work the same way. A personal agent might autonomously reorder toilet paper, while a person buying a jacket may still want to inspect the product and make the final choice through a visual interface.
That means websites will need to support different levels of delegation and involvement, rather than treating “agentic commerce” as a single interaction pattern.
Technologies such as WebMCP could allow a site to expose structured tools directly to an agent, while MCP Apps and other generative interfaces could bring interactive product experiences into the user’s chat environment. An A2A backend might allow agents to interact without traversing the conventional visual site at all.
It might end up being one site with both visual components and agent-accessible tools — two distinct experiences — or perhaps a human-facing website paired with an agent-to-agent service.
“That’s still what everybody’s trying to figure out,” Sanchez said. “But there’s going to be agentic targeting, for sure.”
Whither websites?
Whether websites survive the AI era at all is another big question we’re all grappling with.
What I gleaned from Sanchez at AIEWF was that the traditional website is unlikely to disappear completely, but its role will surely change.
Rather than being a fixed collection of pages that every visitor navigates, a “website” could become a governed content and interaction system that assembles an appropriate interface on demand. At least, that’s the future that Adobe is actively exploring.
Cyberattacks used to be the biggest security issue surrounding AI infrastructure, but that could be changing. A recent cargo theft outside Chicago suggests another vulnerability — and it’s one that has nothing to do with malware or prompt injection.
Just last week, the Cook County Sheriff’s Office recovered two stolen trailers containing roughly $1.3 million in data center equipment and copper wiring, taken from separate shipments originating hundreds of miles away. One trailer held about $300,000 worth of copper wire — reported stolen in Pine Hill, Alabama — destined for data center construction. The other carried roughly $1 million in data center infrastructure equipment, stolen out of Jacksonville, Florida. Both ended up at the same truck yard in Elk Grove Township, outside Chicago.
Viewed in the context of the AI boom, it highlights that the physical supply chain itself is becoming a new target for bad actors.
Viewed in the context of the AI boom, it highlights that the physical supply chain itself is becoming a new target for bad actors.
A new high-value cargo
We’re all familiar with typical bottlenecks like GPU shortages, power constraints and cooling capacity, which have plagued the AI era since its inception. But we forget that building an AI data center requires an enormous volume of specialized hardware moving through freight networks. These include servers, networking gear, fiber, switchgear, cooling systems, power distribution equipment and thousands of pounds of copper. Each represents capital investment and potential deployment delays.
As hyperscalers accelerate the construction of data centers, the exposure of these items between the factory and data center creates a risk category that the industry as largely ignored.
When one delay cascades
Large GPU clusters depend on the synchronized delivery of dozens of interconnected systems. A training cluster is a tightly coupled system of servers, switches, optics, power distribution, and cooling that must be installed together. Missing networking hardware can idle racks, delayed power equipment can postpone an entire deployment and stolen copper can stall electrical work. So when one component category disappears, the delay cascades across everything.
So when one component category disappears, the delay cascades across everything.
Cargo theft by the numbers
Infrastructure resilience increasingly depends on whether critical hardware arrives at the construction site at all — and on schedule. Verisk CargoNet reported that U.S. and Canadian cargo theft losses jumped roughly 60% in 2025 to nearly $725 million, even as the total number of incidents held essentially flat — a sign that thieves are becoming more selective about high-value freight. Metal theft rose 77%, driven largely by demand for copper, while organized groups shifted toward enterprise computing hardware. CargoNet expects that focus on high-value technology — RAM modules, storage drives and enterprise computing equipment — to carry into 2026. For broader context, the Department of Homeland Security has estimated that cargo theft overall costs as much as $35 billion a year.
The Chicago incident fits squarely inside that trend.
Beyond firewalls and malware
Obviously, cargo theft isn’t an engineer’s problem. But organizations building AI infrastructure may need to broaden their thinking about deploying AI capacity on aggressive timelines.
Cloud providers, colocation operators and hardware vendors have already invested heavily in defending infrastructure from digital threats. As AI infrastructure becomes more valuable, protecting the physical systems behind it may deserve similar attention.
The next supply-chain conversation
The AI boom has already forced the industry to rethink electricity, cooling, networking and semiconductor manufacturing. Physical logistics may be next.
It starts long before the equipment reaches the data center.
If the value of AI infrastructure continues to climb into the billions of dollars, the industry’s definition of “infrastructure security” is likely to expand beyond firewalls and identity management. It starts long before the equipment reaches the data center.
Microsoft’s latest AI services announcement suggests the era of standardizing on a single model may be ending. This week, the company launched a $2.5 billion AI adoption business designed to help enterprises customize AI deployments and use multiple models rather than lock themselves into a single provider — part of a larger shift toward systems that route each request to the model best suited to the task.
Simply put, the company that arguably has the deepest single-model partnership in the industry is now selling model swappability as the product.
Betting $2.5 billion on flexibility
Microsoft said Thursday it is creating a new operating entity, Microsoft Frontier Company, to help corporate customers select AI technologies that actually work for their businesses and produce a return on investment, Reuters reported. The unit launches with $2.5 billion in funding from Microsoft and will work with customers including Unilever and Novo Nordisk.
The new firm will help customers choose and integrate AI tools — from Microsoft and external providers — with each customer’s internal data. Customers will own the results of that work rather than handing it back to Microsoft. The move puts Microsoft alongside Palantir, which is doing similar work with large customers using Nvidia’s open-source models, and Amazon Web Services, which recently launched a $1 billion embedded-engineering unit of its own.
What’s most telling is the reasoning. Judson Althoff, CEO of Microsoft Commercial Business, told Reuters the new firm grew partly out of Microsoft’s own experience watching models like DeepSeek and Google’s Gemini catch up to OpenAI. Referring to the original Copilot, he said, “we made a mistake by binding it to OpenAI models only.” Customers, Althoff said, care more about the combination of their data and the models than about any particular model — and they need the ability to swap models quickly as the state of the art shifts.
“We made a mistake by binding it to OpenAI models only.”
One model no longer fits
Consider a typical customer service application that might need to summarize a support ticket, analyze a 300-page contract, generate an email, transcribe a meeting and review source code. Those aren’t necessarily the same problem. A model like Google’s Gemini, with a context window of a million tokens or more, may be the right choice for the contract. A small, fast model like OpenAI’s GPT-5.4 mini or Anthropic’s Claude Haiku may handle ticket summaries at a fraction of the cost. The transcription may go to a purpose-built model like Whisper. And if regulators require customer data to stay on-premises, an open-weight model like Meta’s Llama or Mistral is often the preferred choice.
Instead of choosing a single foundation model, developers increasingly choose several — and the application decides which one handles each request.
AI gateways become core infrastructure
The model is just one component of the stack, so the decision to route the request has to live somewhere. That’s why developers are forgoing hard-coding an application to a single model and building systems that can choose among several. The routing logic might prioritize cost for one request, speed for another, or keep sensitive workloads on a local model. That way, if one provider experiences an outage, traffic can be routed elsewhere without changing the application itself.
The company that arguably has the deepest single-model partnership in the industry is now selling model swappability as the product.
That changes what developers build
Once companies stop relying on a single model, the challenge shifts to building the systems that decide which model to use for each request.
That means developers need tools to route requests, compare model performance, monitor reliability, control costs, enforce security policies and switch to another model if one goes down, which is a very different engineering problem, especially since deciding which model should respond to a request occurs every time someone uses your application. At enterprise scale, those decisions happen millions of times a day, so they have to be fast, reliable and easy to manage.
The ecosystem is already responding
Open-source proxies like LiteLLM and gateways like Portkey normalize APIs across providers. Orchestration frameworks such as LangChain and LangGraph assume the presence of multiple models from the start. The Model Context Protocol (MCP) is making tool integrations portable across models rather than bound to one vendor. And the cloud providers themselves — Amazon Bedrock, Azure AI Foundry, Google Vertex AI — now expose many models behind a single API.
Orchestration is the new moat
Core models will keep improving. But as performance converges for many business tasks, orchestration becomes the challenge. Microsoft’s announcement is one indication that the largest vendors believe enterprises are heading in that direction and are willing to spend billions to be the ones holding the routing layer.
Instead of treating the model as the platform, enterprises are consistently treating it as a replaceable component behind an orchestration layer.
The cloud era taught developers not to tie applications too tightly to one server; containerization made infrastructure portable. Now the same philosophy is being applied to AI. Instead of treating the model as the platform, enterprises are consistently treating it as a replaceable component behind an orchestration layer.
As enterprise AI systems scale to handle complex workflows, practitioners face the challenge of routing subtasks to the right tools and skills. Agents can have hundreds of tools and skills and get confused on which one to use for each step of a workflow.
To address this challenge, researchers at Alibaba developed SkillWeaver, a framework that creates an execution graph for a given task and chooses the right skills for each of the nodes. They also introduce Skill-Aware Decomposition (SAD), a novel technique that uses a feedback loop to enable the agent to fetch and vet relevant tool candidates iteratively. This compositional approach and feedback loop mechanism distinguishes SkillWeaver from other tool-routing frameworks that choose tools in a one-shot fashion.
SkillWeaver relates to real-world AI applications where agents autonomously orchestrate multi-tool ecosystems, such as the Model Context Protocol (MCP), to execute multi-step business operations like downloading datasets, transforming information, and creating visual reports.
In practice, the researchers' experiments with SkillWeaver show that implementing this retrieve-and-route approach significantly increases accuracy while reducing token consumption by over 99% compared to naively exposing agents to an entire tool library.
For practitioners building AI agents, the main takeaway is that the granularity of task decomposition is the biggest bottleneck to accurate tool retrieval.
The challenge of skill routing
Skills are a key pattern in modern LLM agent architectures. A skill is a modular, reusable tool specification that uses structured natural language documentation.
As enterprise agents integrate with massive tool ecosystems, accurately routing user queries to the right skills becomes a difficult task. Exposing an entire library to an LLM to find the right tool is highly inefficient, quickly overwhelms context limits, and consumes hundreds of thousands of tokens.
Most current tool-use frameworks attempt to solve this through API retrieval, documentation matching, or hierarchical structures that treat routing strictly as a single-skill selection or per-step problem.
However, this single-skill paradigm is insufficient for enterprise environments because real-world queries are inherently compositional. A standard business request such as "Download the dataset, transform it, and create visual reports" cannot be fulfilled by one tool. It requires breaking the prompt down and sequencing an API client, a data processor, and a visualization tool into a cohesive, multi-step execution plan.
How SkillWeaver and SAD work
To tackle this, the researchers frame the problem of handling complex tasks that require multiple skills as "compositional skill routing." Given a complex user prompt and a vast library of tools, an agent must simultaneously figure out how to break the request into a sequence of atomic sub-tasks, how to map each sub-task to the single best available skill, and how to compose those skills into an executable plan.
SkillWeaver orchestrates this process through three distinct stages: Decompose, Retrieve, and Compose. In the first stage, an LLM acts as a task decomposer, breaking the user's complex query down into a sequence of sub-tasks that each require one skill. Once the sub-tasks are clearly defined, the system uses an embedding model to compare each subtask against the skill library to pull a shortlist of the top candidate tools for each step.
In the final stage, a planner evaluates the retrieved candidates based on how well they work together. It checks for inter-skill compatibility to ensure the outputs of one tool naturally flow into the inputs of the next. It then creates a final execution plan as a Directed Acyclic Graph (DAG) that maps out dependencies so independent tasks can potentially execute in parallel.
For example, consider a user asking an AI agent to "Download the dataset, transform it, and create visual reports." In the decompose stage, the decomposer LLM breaks this into three distinct sub-tasks: downloading the dataset, transforming the data, and creating the reports.
In the retrieve stage, the system searches the library and finds candidates like “api-client” or “http-fetch” for task one, “csv-parser” or “etl-pipeline” for task two, and so on. Finally, the compose stage evaluates these options, selects the specific combination of “api-client,” “csv-parser,” and “chart-gen” that are most compatible, and wires them together into a final, ready-to-execute workflow.
A key challenge of this pipeline is that LLMs often produce generic step descriptions that fail to match the specific, technical vocabulary of the actual skills available in the library. To fix this, SkillWeaver introduces Iterative Skill-Aware Decomposition (SAD), a novel feedback loop. SAD works by having the LLM draft an initial plan, conducting a preliminary search to find loosely matching skills, and then feeding those retrieved skills back into the LLM as hints. This allows the LLM to rewrite its decomposition so the granularity and vocabulary perfectly align with the actual tools that exist.
SkillWeaver in action
To evaluate how SkillWeaver performs in realistic enterprise scenarios, the researchers created a custom benchmark called CompSkillBench. It consists of 300 multi-step queries of different difficulty levels. To mirror real-world environments, they used a library of 2,209 real-world skills sourced from the public MCP ecosystem, covering 24 functional categories like cloud infrastructure, finance, and databases.
For the core engine, the researchers primarily used a lightweight 7-billion parameter model (Qwen2.5-7B-Instruct) for task decomposition, paired with a standard semantic search retriever (MiniLM with a FAISS index) to find the tools. SkillWeaver was evaluated against three main setups: a brute-force "LLM-Direct" method where they stuffed all the tool names into the prompt of a large model, a vanilla LLM-based decomposition without SAD, and a ReAct-style agent loop.
The experiments indicate that task decomposition is the main bottleneck. Standard LLM behavior falls short when dealing with large tool libraries, but the SAD feedback loop dramatically moves the needle. In the vanilla setup, the 7B model achieved a decomposition accuracy (i.e., predicting the correct number of steps) only 51.0% of the time. By activating the SAD feedback loop, accuracy jumped to 67.7% (with the larger Qwen-Max model, the accuracy reached 92%). On "hard" tasks requiring four to five distinct skills, SAD improved accuracy by 50%.
One fascinating finding was that larger models can actually perform worse when unguided. When tested in the vanilla setup, a larger 14-billion parameter model saw its accuracy plummet below the 7B model's accuracy because it tended to over-decompose tasks into microscopic, unnecessary steps. Once SAD was introduced, the retrieved tool hints anchored the model back to reality and increased its accuracy. This suggests that aligning an agent with the vocabulary of specific tools is often more impactful than paying for a larger, more expensive LLM.
Another important takeaway is token savings. The LLM-Direct baseline, which used the very large Qwen-Max model, showed that feeding all tools into the prompt of a large model fails. Despite near-perfect task breakdown capabilities, the massive model only retrieved the right tool category 21.1% of the time when flooded with tool options. SkillWeaver's targeted retrieve-and-route approach vastly outperformed this in accuracy while slashing context window consumption from an estimated 884,000 tokens down to roughly 1,160 tokens per query, a 99.9% reduction. For practitioners, this translates directly to drastically lower API costs and faster response times.
Finally, the traditional ReAct baseline completely failed, achieving 0% decomposition accuracy. Its loop naturally collapses multi-step plans into isolated actions rather than explicitly mapping out a cohesive, multi-tool sequence.
Considerations for developers
While the researchers have not yet released the source code for SkillWeaver, their work was built on off-the-shelf tools that can easily be reproduced.
Skill-Aware Decomposition (SAD), which is the key innovation at the heart of the framework, is a clever prompt-engineering and retrieval loop. The authors have shared the prompt templates in their paper, and developers can implement it themselves quite easily using standard orchestration libraries like LangChain, LlamaIndex, or even raw Python scripts.
As for the retrieval component, the authors built the core framework using all-MiniLM-L6-v2, an open-source embedding model. They found that swapping in a slightly stronger off-the-shelf encoder (BGE-base-en-v1.5) immediately boosted accuracy without any fine-tuning. While an off-the-shelf bi-encoder is great at getting a relevant tool into the top 10 candidates nearly 70% of the time, it struggles to consistently rank the perfect tool at exactly number one, achieving that only about 37% of the time. To bridge this gap, teams will likely need to implement a secondary cross-encoder or LLM-based reranker to re-order those top 10 candidates.
One upfront preparation requirement is vectorizing the tool library and building a FAISS index in advance. In practice, this is a negligible hurdle. Embedding and indexing all 2,209 skills in the benchmark took a mere 15 seconds. Once built, retrieving tools from the index adds less than 15 milliseconds of latency per query. For enterprise environments, syncing the tool index is a trivial background job.
A current limitation in SkillWeaver is the lack of error recovery. While SkillWeaver successfully maps out a compatible DAG for execution, the authors' pilot study revealed the challenges of multi-step tool chains. For example, if an API call fails in step two, the entire chain breaks. The paper's core contribution is limited to the routing and planning phase. For a true production deployment, practitioners must build their own error recovery, fallback, and retry mechanisms on top of the compose stage to handle real-world API timeouts or malformed outputs.
Researchers have found a never-before-seen piece of macOS malware that combines a series of clever tradecraft to infect Macs with stealthy, custom-developed credential-stealing code.
The malware is delivered in two stages. The first is distributed in a disk image that masquerades as Maccy, a clipboard manager for Macs. It’s compiled as AppleScript that is notable for the way it delivers the second stage. The malware is named PamStealer because the Rust-written infostealer uses the Pluggable Authentication Modules interface built into macOS to validate the target’s login password before sending it to an attacker-controlled server.
A quieter execution chain
The use of both disk image and AppleScript is common in malware for Macs. More unusual is the way PamStealer combines them to gain stealth. When the AppleScript is double-clicked, it’s opened in the macOS Script Editor, where the malicious functionality is buried deep within the file.
Enterprises are expanding their use of automation in IT, where AI is changing the landscape with trends such as agentic workflows, automated decision-making and AI-driven pipelines.
When Subquadratic launched earlier this year, it could build a sparse-attention model that could handle a 12-million token context window and be significantly faster than today’s large language models. But it didn’t launch the model widely and it didn’t publish benchmarks.
Given the company’s large claims, that created quite a bit of skepticism. In June, Subquadratic published its first model card and benchmarks for its small model, SubQ 1.1, supplied third-party verification from data firm Appen, and started talking about its first design partners who now have access to its model.
So far, however, few people have actually used its model. To talk about the company, why its model isn’t widely available yet, and what it has in store for the near future, we met up with Subquadratic co-founder and CTO Alex Whedon.
“We’re not a sparse attention company either.” — Alex Whedon, Subquadratic.
One thing Whedon definitely wanted to clear up is that the company’s current model may be based on sparse attention, but that isn’t its full mission.
“We’re not a sparse attention company either,” Whedon tells The New Stack. “We’ve been working on non-attention architectures for quite a while as well. We think that we will be the first people to leapfrog ourselves in terms of the next model architecture.”
We’ll get back to that.
What the model card shows
It’s the company’s SubQ 1.1 Small model that people are talking about now. This model is built on Subquadratic Sparse Attention (SSA), an attention mechanism the company says scales close to linearly with context length instead of quadratically.
“In the case of Subquadratic Sparse Attention specifically, which is one of a couple model architectures we worked with, the idea is that not all of the token relationships matter,” Whedon explains. “Token relationship compute is why you see this quadratic scaling law.” This means there are almost a million possible two-token relationships in a 1,000-token input in a full attention matrix.
For SubQ 1.1 Small, the strongest results are in long-context retrieval, which makes sense, given that this is where the architecture should have its biggest edge.
Credit: Subquadratic.
On the needle-in-a-haystack test, SubQ 1.1 Small scores near-perfect from 1 million tokens out to 12 million, even though it was trained mostly at 1 million. It hits 99.12 percent on Nvidia’s harder RULER test, which asks the model to trace and aggregate facts across a 128,000-token context rather than just find one.
On general capability, it lands just below the mid-tier frontier models, at 85.4 on GPQA Diamond against 87.5 for Sonnet 4.6. On the LiveCodeBench coding benchmark, it scores 89.7, below Opus 4.8 and GPT-5.5, but slightly better than Sonnet 4.6.
Efficiency is where the model shines, though. The company says that at 1 million tokens, SubQ uses 64.5x less compute than dense attention and runs 56x faster than FlashAttention-2 on a single attention layer. At the full 12-million-token window, it puts the attention compute reduction at close to 1,000x.
Credit: Subquadratic.
“Even in full dense attention, the relative importance of over 99 percent of tokens is very low, attention scores are below 0.1,” Whedon says. “We actually show this in our model card. So clearly we’re just wasting compute most of the time, and in fact we’re maybe making the modeling task harder, because we’re introducing noise.”
“Transformers are a brute-force approach to the problem of text modeling,” he says. “You could say, ‘I’m going to compare every single individual token to every other possible individual token.’ That’s what transformers do. Very brute force, very naive. It just assumes that the first needs to look at the second, the third, the 50th, and the 5,000th. That’s not how humans read text.”
SSA also differs from retrieval-augmented generation, which drops chunks of text before the model sees them. “Every token of the text is being seen by the model,” he says. “It’s just not being redundantly compared to every other token of the text.”
On capability, SubQ 1.1 Small lands roughly in Sonnet 4.6 territory, sometimes a bit above, sometimes below. But its edge, the company says, is size and cost.
“What we posted publicly was fewer than 100 billion parameters,” Whedon says about the size of the model. “I would venture to say that our model is smaller than any of the models offered by OpenAI or Anthropic. But our next model will not be.”
Smaller, cheaper, built for enterprises
Subquadratic is also making the pitch that its model’s capabilities will be especially interesting for enterprises.
“We think that’s a pretty interesting enterprise offering,” he says. “We’ve seen a lot of people in the enterprise space talking about using the mid-tier models as opposed to the frontier for large data-processing tasks, which is exactly where we’re trying to plug in.”
Given that a lot of enterprise problems start with searching through large heaps of data, this makes sense. You can pack a lot of documents into a 12-million token context window, after all. Most of today’s models break down well before the user fills their million-token windows, but with its near-perfect retrieval scores, SubQ may be a good answer for these problems.
As Whedon noted, the model’s first users are design partners, not the public. “We’re giving access to the model to design partners now, and these are mostly enterprises, largely with eight- to nine-figure spend,” Whedon says. “This is a core market that we really care about. It has been since day one.” A limited individual-access release will follow before any general availability.
The launch led with claims instead of benchmarks by choice.
“We were announcing mostly research,” he says. “We could have maybe messaged the launch a little bit differently. There was some debate about how we were going to message it.”
Built on an existing model
One question from May hasn’t gone away, though. The model card states that Subquadratic “started with an existing open-weight frontier model by replacing its dense attention with Subquadratic Sparse Attention (SSA),” and then ran roughly one trillion tokens of long-context continued pretraining on books, documents, and repository-scale code.
That confirms what some of the skeptics suspected at launch, when OpenAI researcher Will Depue wrote that SubQ was “almost surely a sparse attention finetune of Kimi or DeepSeek.” What’s new here then is the SSA mechanism and the long-context training recipe, not a model trained from scratch. The company has not said which open-weight model it started from.
The biggest lever on long-context retrieval was pretraining on very long sequences, Whedon says, something SSA’s efficiency made cheap enough to run as routine.
“Nobody’s talking about multimillion-token pretraining,” he says.
Credit: Subquadratic.
Why hybrids don’t go far enough
There have, of course, been attempts to improve on quadratic scaling, but Whedon thinks most of those attempts only go — almost literally — halfway. Hybrid models such as Nvidia’s Mamba-based Nemotrons, Qwen’s Gated DeltaNet layers, and the various linear-retention designs swap out some of the attention layers, but they don’t go all the way.
“If 80 percent of the layers are not quadratically scaling, then your maximum payoff is like a 5x increase as you scale toward infinity,” he says. “We see a 60x increase at 1 million tokens, almost 1,000x at 12 million. That is the type of payout that you only get if you actually change the scaling law, as opposed to a scalar win.”
He actually credits DeepSeek’s own sparse attention mechanism with making his company’s pitch easier.
Credit: Subquadratic.
“They showed that you could dynamically select relationships without a significant quality trade-off,” Whedon says. “However, they did so by redundantly using a smaller but still full-attention model that ends up using the vast majority of the compute at scale.”
Subquadratic ran its own benchmark against GLM 5.2. “At 1 million tokens, 58 percent of the prefill latency comes from that selection mechanism,” Whedon says. “So that selection mechanism, which is supposed to be seen as cheap, actually dominates the compute, because it’s a quadratically scaling component.”
Beyond sparse attention
It’s also why Whedon pushes back on the “sparse attention company” label. Subquadratic has been working on what he calls “zero attention,” architectures that drop the attention mechanism altogether.
“Attention is kind of similar to RAG in that you have queries, keys, and values that represent information about the tokens that you’re processing,” Whedon says. “There’s this discreteness of representation, where everything is represented within these nice little boxes. That’s super convenient. It’s easy to build a brute-force solution around it. But it also means your ability to compress information is limited. If you had a more continuous, abstract way of representing the information, then you could compress it further, which means you can make smaller models, or you could just scale things up again to create another leap in intelligence.”
He traces the idea to world models and to Yann LeCun’s work. “The stuff we’re doing takes a lot of inspiration from world models, not the video modality in this case, but some of the things LeCun is talking about,” he says. “Rethinking how to represent long-range dependencies, how to keep a long-range state, how to rethink the objective function.” He stops there. “That’s probably all I could say for now.”
Subquadratic has also marketed only one of the three kinds of efficiency it says it is chasing. “We care about compute, sample, and memory efficiency,” Whedon says. “We’ve done a lot of work on all three, but have only really talked about the compute efficiency publicly.”
The near-term plan
The near term plan for Subquadratic, however, is more modest. “Over time, yes,” Whedon says, when asked whether Subquadratic could rival OpenAI and Anthropic on raw quality in the long run. “In the shorter term, we have to be strategic. If we try to boil the ocean on much less capital, it’s not going to go well for us.”
The next model, he says, will likely be a mid-tier size rather than a frontier-class one (think SubQ 1.2 Medium), that he expects to outperform most of the competition in its tier.
How the team will bring the model to market, though, remains to be seen. I wouldn’t be surprised if the team launched its model on one of the hyperscaler’s large model platforms, but Whedon remained tight-lipped about the company’s plans.
The fact that we met with the Miami-based Whedon in San Francisco, though, gives you a bit of a hint of what the team is currently up to.
New capabilities, such as data quality agents and a feature that makes data products more reusable, support engineers to help organizations more easily achieve their AI goals.
Frameworks like Lean Six Sigma and business process management (BPM) first gained traction because they promised clarity in the chaos—a structured way to bring order to messy, sprawling operations. Lean Six Sigma emphasized statistical rigor and quality control; BPM created end-to-end maps of how work should flow across departments. Both offered a repeatable way to embed habits of measurement, analysis, and accountability into day-to-day company culture.
But today, those time-tested playbooks are evolving as companies seek to embed AI into established process excellence methodologies. By some estimates, the market for AI-powered process optimization is projected to exceed $113 billion within the next decade. In one study, a full 88% of business leaders anticipated increasing investments into AI-infused process intelligence in the next 12 to 18 months.
Yet without the right foundations, many of those investments may not fully deliver on their potential. Companies that already operate with discipline have an edge. They can channel new tools into proven systems rather than bolting them onto shaky foundations. Organizations with mature process disciplines are also better positioned to translate AI ambition into real outcomes, as they are already accustomed to data-driven decision-making and process discipline—precisely the cultural foundation AI systems need to deliver value.
Simply put: AI can accelerate process excellence, but existing process excellence is what makes AI truly impactful. Technology and process are no longer separate levers, and only organizations that pull them together stand to realize the full value of both.
This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.
When building ML models, developers can use several techniques to make models easier for humans to interpret, leading to improved transparency, troubleshooting and user acceptance.
Impeccable’s Paul Bakaus at the AI Engineer World’s Fair.
Paul Bakaus thinks the emerging discipline of “skill engineering” can make AI agents more capable — but he absolutely does not want to remove people from the creative process. He chats to Latent Space about his approach to design in the AI age.
Bakaus is the creator of Impeccable, an open-source design skills system that gives coding agents a vocabulary for improving interfaces. Instead of asking an agent to redesign an entire website in one shot, users can tell it to make a section “bolder,” “quieter,” “denser,” or more polished.
Behind those apparently simple commands is a larger argument about how AI products should be built. Agents need more than instructions, Bakaus said: they need domain knowledge, context and carefully defined ways for humans to steer the result.
“The point is to give you a way to steer what you want to end up with,” he said during a session at the AI Engineer World’s Fair. “It’s never going to be a tool for one-shot design. That’s not the intent.”
The emerging craft of skill engineering
Impeccable began as a relatively simple extension of Anthropic’s frontend design skill. As its audience grew, Bakaus expanded it into a more complex system with multiple components and workflows.
That process led him to start thinking of skill engineering as a discipline in its own right. His workshop at the conference explored what he called the “dark arts” of building skills.
“One of the interesting topics was that most skills — [and] most models — are not very creative,” Bakaus told me. “They converge in one direction, and if everybody uses the same skill to do frontend design work or something like that, everything ends up looking the same.”
Skill engineers must also account for differences between agent harnesses and models. Codex and Claude, for example, do not necessarily handle subagents or permissions in the same way. A skill intended to run across Claude Code, Cursor, GitHub Copilot and Codex cannot assume they all provide identical capabilities.
Bakaus has also experimented with routing inside a skill, allowing it to combine several capabilities and direct a task toward the relevant instructions. He compared this to a mixture-of-experts model, with routing used both to conserve tokens and improve effectiveness.
Giving agents a design vocabulary
Impeccable’s core innovation is to take terms familiar to designers and give them a more precise operational meaning for an agent.
An unassisted model asked to make a page “bolder” may add gradients, neon effects or glass-like surfaces. Impeccable instead defines boldness through concepts such as hierarchy, scale and decisive typography — changes that attract attention without necessarily breaking the existing design system.
“An adjective with nothing behind it is just a nice apostrophe,” Bakaus said. “You really have to tell the agent what you mean.”
He described these terms as words that have been “imbued with meaning.” The model already has some conception of what words such as “bold” or “quiet” mean, but the skill translates them into a specific professional domain.
This is the key, because experts often possess a vocabulary that non-experts do not. Bakaus said he had observed large differences between the work produced by a designer and an engineer using the same model, simply because the designer knew how to articulate the desired result.
“I’ve been trying to put that language — basically compress it into a skill and into a system — to be able to express yourselves better,” he said.
However, he does not believe every part of design can be controlled from this level of abstraction. Directly manipulating spacing may still be the fastest option for a small adjustment, while open-ended prompting can be useful during initial exploration.
The objective is not to replace every tool with an agent, he insisted. It is to determine “the exact level of control” and insert the person at the point where their judgment is most valuable.
Designers and engineers move up the stack
Bakaus sees the boundaries between design, engineering and product management becoming less distinct.
“Designers are moving into code, engineers are moving into design, and vice versa,” he said. “These worlds are all colliding.”
That shift will be uncomfortable for people whose work primarily consists of translating an existing artifact into another form. Engineers who mainly turn Figma designs into code face growing automation, while designers whose contribution is limited to making an existing interface look competent face similar pressure.
“Designers all have to move one layer up the stack to think more about the what,” he said. “I think the role of the product manager and designer is actually converging.”
At the same time, designers are moving closer to implementation — into code. Bakaus initially expected Impeccable to appeal mostly to engineers and assumed professional designers might resent that. Instead, he estimates that designers now make up at least half of its audience.
“So rather than moving directly into code and, you know, having no help,” Bakaus said about designers, “they use Impeccable as a bridge, because it communicates the way they communicate. And that was not obvious to me when I first built it.”
Impeccable also has a live mode that combines visual selection with an underlying coding agent. A user can select a section inside a development environment and request several alternative layouts or (for example) ask for a bolder or quieter treatment. The system operates within the project’s existing code and design system rather than exporting an isolated mockup from a third-party design tool.
Bakaus described this as a potential “design harness” at the intersection of chat and direct visual manipulation.
There will be no auto mode
The AI industry often treats complete automation as the natural endpoint of product development. Bakaus rejects that premise.
He sees two dominant camps: people trying to preserve the traditional Figma-centered workflow, and on the other side advocates of “loopmaxxing” who want agents to work with as little human intervention as possible.
“The truth is somewhere in the middle,” he said.
His preferred model is for AI to produce the first 80% quickly: the competent layout and basic implementation that would otherwise consume a lot of time. The person then owns the final 20%, where taste, context and a distinctive point of view enter the product. This is a key part of Bakaus’s design philosophy in the agentic era.
“People need purpose, and they want to play a role in whatever they create,” Bakaus said. “When you work with the agent, then you feel more ownership of the product.”
Users regularly ask him to add an automatic mode to Impeccable so that the system chooses the commands itself. He has no intention of doing so.
“There is no auto,” he said, “and there will be no auto.”
Asked about the language of software factories and other visions that appear to remove people from engineering altogether, his response was unambiguous.