❌

Reading view

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time.

One new feature in particular caught our eye: WorldPrompt, a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events. The events, or actions, can even be prompted in real-time.

To understand the implications of WorldPrompt, we spoke to Kamil Sindi, Runway’s CTO, and Robin Kahlow, its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis, co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him.

Who’s building real-time interactive world models?

First, some context about world models that can generate interactive video and audio in real-time.

Runway is reportedly valued at $5.3 billion, based on its most recent fund raise of $315 million in February. Its first release, GWM Worlds, was launched last December.

Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro, and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table:

Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations. For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.”

But as our interviews with Runway show, real progress is being made.

The central idea of WorldPrompt

WorldPrompt, a new feature in GWM Worlds 2, helps differentiate Runway from its competition. You can think of it as a control layer for characters, cameras and the environment. As Kahlow put it, it’s a way to “control all the different subjects in the world” — similar to a computer game.

“Like, if there’s an NPC [Non-Player Character] somewhere, the NPC might walk up to you and say something. So you could achieve the same thing with this kind of model, where you can have very detailed control over everything in the scene.”

As the name suggests, WorldPrompt is a prompting mechanism — not a programming language. So, unlike virtual world games like Minecraft or Roblox, GWM Worlds 2 doesn’t offer scripting capabilities or the ability to control state. But there’s a power to that, as Sindi pointed out.

“You can create promptable worlds on-demand with video and audio in sync, across all these different domains and environments. That’s not a distant-future hypothetical thing,” he said.

But there are also limitations to prompting a world model. We asked how reliably the model would follow an instruction to create, for example, a law of gravity or a certain ability in a character?

“Yeah, so it’s a research preview,” Kahlow replied. “So it’s not perfect, of course, and there are still flaws. It really depends on how difficult the action is. I would say movement works quite reliably.”

Sindi added that more training plus scaling the data and models is resulting in “better following.”

How a video model becomes a real-time runtime

Despite the current limitations of GWM Worlds 2 — especially if you compare it to pre-designed and scriptable worlds like Minecraft or Roblox — the true promise of world models like Runway is that they’ll eventually lead to fully self-generated, real-time games and experiences. Which is an extremely hard engineering problem, as Kahlow reminded us.

“There are two challenges. One is making the model not generate a whole clip at once. So instead, you want it to generate frame by frame while you’re looking at it. And the other challenge is actually making the generation fast, so you can play it in real time.”

High-level view of GWM Worlds 2 process

GWM Worlds 2 offers real-time interactive worlds streamed in continuous 720p video at 24 frames per second (fps) and audio at 48,000 Hz.

Runway achieved this firstly by taking its foundational audio-video generation model and fine-tuning it to the new WorldPrompt format, so the model can follow that. It then post-trains the model to generate autoregressively.

“And after that, we work on making it real-time through distillation methods,” Kahlow added.

Co-CEO Anastasis Germanidis offered more technical details in our podcast with him. He told us that the process starts from “bidirectional diffusion that basically generates an entire video at once and [makes] it autoregressive.” This allows the model to “generate one frame or a few frames at a time.”

Autoregressive causal diffusion vs traditional video models; image via Runway

Germanidis described two possible forms of distillation in order to make it real-time: distilling a larger model into a smaller one or reducing its diffusion steps. As a general example, he said a model might go from around 50 denoising steps to four, with some quality loss but potentially comparable results.

The challenges of real-time generation

Germanidis admitted that there were issues with how it generates real-time interactive video.

“The biggest challenge with autoregressive models is error accumulation,” he said. “You’re feeding generated frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time.”

Errors compound; image via Runway

Sindi told us there are also challenges dealing with “infinite generations” of content.

“There’s all these challenges around what context to keep, what to discard that’s not important. And so there’s all these optimizations we have to think about, so we’re not blowing up our GPU memory.”

Another current limitation is long-term memory. “The model does not have perfect memory,” Kahlow said. “That’s still an open research problem.”

Causality and correctness

While performance is the primary challenge for Runway at this time, its world model also has to produce plausible consequences when a user takes different actions.

Germanidis used the example of simulating football; he pointed out that online video training data contains more successful goals than failed goal attempts, so a video model might render the first more convincingly.

“If I take this action versus this action, you want it to generate equally realistic outcomes,” he told us. “That’s, I think, the big gap between video models and world models: that idea of counterfactual generation.”

Sindi told us that evaluation gets harder the more complex interactions get.

“If you have this multi-prompt, multi-character, multi-scene [environment], how do you really understand what was causal and what was not?”

To try and solve that, Runway has some automated verifiable tests. But since GWM Worlds 2 is a research preview, Kahlow noted that doing tests yourself is also advisable — “trying out your model to see what doesn’t work is really important.”

More than gaming — there are agent use cases too

Gaming is the obvious use case for what Runway is building, but there are others. Kahlow mentioned robotics — for example using a simulated environment to test how a robot works.

Another, more intriguing, use case is to use it to test agents at scale.

“Having thousands of simulated environments is much less challenging if you have a suitable model like GWM Worlds,” Kahlow said.

But how does an agent know what’s changed in the world — is there a structured state that it can read, or is it just the generated video and audio that it’s consuming and understanding?

“So there’s no structured state here,” Kahlow replied. “It’s just observing the same thing you might observe in real life, just [in this case] from cameras.”

Sindi noted that GWM Worlds can also be used for “synthetic data generation for agents.”

Finally, Germanidis suggested there’s potential to use these world models alongside reasoning models.

“You’re maybe using some reasoning [for] planning of the scene, and then you’re passing it into the diffusion head that’s actually generating the pixels.”


Anastasis Germanidis


Timestamps

00:00:00 Introduction

00:05:17 Runway’s Origins and the Bet on Generative Video

00:12:23 The Stable Diffusion Story

00:18:44 Gen-2, Controllability, and the Weekend Hack

00:23:02 From Video Generation to World Models

00:28:03 Learning From the World, Not Just Language

00:35:04 Sora, Runway’s Existential Crisis, and Gen-3

00:39:39 Why Real-Time Video Is Inevitable

00:43:06 Interface World Models: Software Without Code

00:50:25 The Fully Neural Operating System

00:55:11 World Models for Robotics

01:02:32 Robot Policies and World Action Models

01:07:47 The Lucid Dream Test

01:11:41 Video Agents and Omni Models

01:23:12 Artists, AI, and Creative Workflows

01:27:14 Physical AI and the Future of World Models


Transcript

Introduction: Runway, Creative AI, and the Early Thesis

Swyx [00:00:00]: Okay, we’re here with, Anastassios from Runway, with, me and Vibhu in the studio. Welcome.

Anastasis [00:00:08]: Good to be here.

Swyx [00:00:09]: Congrats on all your success and progress with Runway. You’re opening offices all over the world. Did you envision this when you first started out?

Anastasis [00:00:16]: Not quite. I think even when we started, we had this idea that, It was more a matter of when, not if, we were seeing the early generative models of 2016, 2017, and just extrapolating, assuming, we resolution, quality increases predictably over time. There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of those generative models, rethink how creative tools are made. and as we built out the research behind, our generative models, it then became clear that they were useful far beyond that as well.

Anastasis’ Background: Art, Simulation, and Machine Learning

Swyx [00:00:57]: And it is more obvious now with, like, the real-world stuff and the world models that we’ll talk about later. I’m just kinda curious how you go from a background in, like, Zocdoc and, computer vision into Runway. Like, take us back to that early conversations with Chris and, whoever else is on your founding team.

Anastasis [00:01:14]: I was always splitting through those two worlds. One was the I had my own art practice. I was making a lot of interactive art, I think for a long time. and then on the other side, I was working in startups, and I was working as a ML engineer, as a backend engineer at different companies. I’ve always been interested in, coding and computation, and especially interested in simulation and brought it back into my early artwork as well. And at the same time, I was interested in

Swyx [00:01:43]: The personal site has a few, right?

Anastasis [00:01:44]: Yeah.

Swyx [00:01:45]: Is there one that we should pull up? Just in case there’s something that’s like. I just like to go down memory lane.

Anastasis [00:01:50]: Yeah.

Swyx [00:01:50]: Okay, what is this?

Anastasis [00:01:51]: So this was, a project that I made, I think back in 2015, where I built this software that would give, voice instructions to people in a gallery space. So it would coordinate interactions between people. And so it will first give you an identity, like you’re an, architect, you’re 30 years old, and, you like sports. and then it would match you with another person, and you have this completely generated interaction. language models were not quite there at the time, and so it was it was a mix of some templates and some, like, some Markov chain-generated text, and it would just completely simulate these small talk conversations between, everyone in the gallery space. so was always very fascinated on the one hand with, generative models and, like, the early machine learning work that was being at that time. But at the same time, there was this separate thread of simulation and what it means. Like, what can we learn about humans by creating those very simple models of their interactions and their behavior?

Early Generative Art: pix2pix, GANs, and Uncanny Valley

Vibhu [00:02:56]: Did you generate the prompts or, the 30-year-old, whatever? Was it you generating them? How’d you, how’d you

Anastasis [00:03:03]: Exactly. So the program would just generate- those, from. Yeah, a lot of it would be Mad Libs style of just

Vibhu [00:03:10]: Yes

Anastasis [00:03:10]: You have lists of different professions, lists of different,

Vibhu [00:03:14]: Hobbies

Anastasis [00:03:15]: Personality types, lists of different, ages, things like that. And then it would just combine those things together. And then maybe the next project we go is, Uncanny Valley, Uncanny Road, which was

Swyx [00:03:27]: Gans

Anastasis [00:03:27]: One of the first projects that, we built with, one of my two co-founders, Chris. This was taking, pix2pixHD, which was one of the early image-to-image models that NVIDIA released back in 2016 or 2017. and it was a model that would take a semantic map of a scene and then generate a photorealistic, let’s call it, output. very early days, so it was not very high-fidelity outputs, but it w I think was the first image-generation model that could generate at 1K resolution. And it was all trained on self-driving datasets. So the semantic categories it would support were only, things you would encounter on the road. So it would be pedestrians, traffic signs,

Vibhu [00:04:16]: Stoplights

Anastasis [00:04:17]: Bikes, stoplights. And so that was one of our first indications that we built this and people were making all this, like, very surreal imagery of, yeah, a million plus a million pedestrians or a million traffic signs or, like, gigantic humans. And it was a indication that you could take a model that was trained on this very boring dataset, essentially, of, like, not that many interesting things happen when you’re on the road, and then you can repurpose it and go very out of distribution and make something that was artistically compelling. And that was It’s a summary of the thesis of Runway in some ways, that you can take the same generative models, and if you look at them from another direction, if you build interesting tools around them and you give them to artists, they’re gonna do things that you don’t expect.

Vibhu [00:05:02]: Very cool. I like the, UX of it. You’re just given an empty canvas, try whatever, do whatever. And then the other one, like, you see everyone with wired headphones? Like, that’s, that’s a sign that it’s, it’s very

Anastasis [00:05:16]: The Apple

Vibhu [00:05:17]: Yeah

Anastasis [00:05:17]: Apple, your version.

Vibhu [00:05:17]: Original ads. Yeah. Take us to today. You’ve been doing this for seven years at Runway. How have we got to this? Like, how do we go from driving simulator data to all this? And you cover the whole stack of generative media?

From Creative Tools to a Research Lab

Anastasis [00:05:33]: Interestingly, we’re almost back in, we’re, we’re full circle. We’re, we’re now applying our models and beyond creative tools into real-world scenarios. But it was a, it was a long journey. It was very early on we realized the first version of Runway was a way to easily use the, all the open source model of the day, things like pix2pix to. and give them to artists. That was the initial idea, is those models are too difficult to use if you’re not a machine learning engineer. Like, what happens when you give them to artists? Very quickly, we realized we needed to build a research org, inside of Runway, and that happened maybe on year one. And, a lot of the mandate there was. The image-generation models of the time, the video generation models of the time, or there were barely any video generations all the time, but they were not quite there where they could be productionized and brought into tools that would be part of creative workflows. so we need to push the frontier of the research. And so maybe the first four years of Runway, research was almost happening on the background until there was a moment in 2022, with latent diffusion, with, DALL-E 2, where, there was that step function change, and you guys maybe remember around the time.

Swyx [00:06:49]: I started in this space because of latent diffusion and Stable Diffusion.

Anastasis [00:06:54]: Yeah.

Swyx [00:06:54]: Because I was like, “Wow, this is not only, like, feasible, it is doable on consumer hardware.”

Anastasis [00:07:01]: Exactly, yeah.

Vibhu [00:07:01]: I think the delta is also huge. Like, I learned pix2pix. Like, this was intro to ML, the TensorFlow, like, Jupyter, Google Colab notebooks were like this, and then you have a sudden step function change, with diffusion and whatnot. Any other ones since that. Like, there were clear examples of what early diffusion were to get to here. Any other changes in key technology research?

Green Screen, Rotoscoping, and Early Runway

Anastasis [00:07:26]: Between, 2018 when we started and 2022?

Vibhu [00:07:29]: Yeah.

Anastasis [00:07:29]: So one of the early work that we did in Runway was solving segmentation, image and video segmentation. It was a very important problem because most VFX involves essentially separating

Swyx [00:07:42]: Rotoscope

Anastasis [00:07:42]: Subjects. Yeah, rotoscoping. Extremely manual process. Nobody enjoys doing that. and so a lot of the early days of Runway was building this tool. It was called Green Screen, and it was for a long time the main thing that people were using Runway for. It ended up being used in, Everything Everywhere All at Once and a bunch of other high-visibility films and series. But that was essentially, Runway for a long time was a post-production tool until latent diffusion and generat- Gen-1, Gen-2, happened.

Swyx [00:08:12]: Cool. let’s, let’s go past that moment. You’ve come a long way. Then you started releasing your own models. Maybe describe that journey as well.

Scaling Video Models and the Bet on 1,000 A100s

Anastasis [00:08:20]: Yeah, so we go to the other point, yeah, in mid-2022 when it became clear that we’re doing research at a fairly small scale of compute, and it became clear that, like, scaling laws would apply to, image and video gen in the same way that we’re applying to language generation. So we made a big bet, and I think at so at the time, we signed this deal to build a cluster of a thousand A100s, which at the time we were a Series B startup. That was a almost, slightly irrational decision maybe, but we really believed that if we trained a video model at a large scale, we would get, like, a great model at the end. And at the time, the goal or we set the goal around fall of 2022 of what is, what does the latent diffusion, Stable Diffusion moment look like for video? And at the time, the best model of the time was called CogVideo. it was one of the early video models. It was very 256 by 256 resolution, very not very high quality. and so we decided we’re gonna build out this cluster, and we’re gonna just invest in, like, in building out our own video model. it became clear as we’re training Gen-1 that it was difficult to get to fully. we wanted to build text-to-video, but it became clear to us that an easier starting point would be to start from video to video. Because when you have a stronger conditioning, it’s, it’s an easier problem to restylize an existing video versus generate the video from scratch. And so we released Gen-1 first back in, it was January of, 2023. Yeah.

Vibhu [00:10:04]: It’s just a fun visual podcast, honestly. Like, if we can see February 2023, what was the state of stuff?

Gen-1: Video-to-Video and Depth Conditioning

Anastasis [00:10:10]: It’s so interesting ‘cause at the time when you see those results, you think this is so incredible, and this is like, it’s almost like image generation or video generation is solved. And then you look back a few years after, and it’s like, it’s It’s just like you get used to the results very quickly, with those models. But at the time when we started seeing those results, it was, it felt quite incredible, and the level of, like, quality that you could get. And, so the Gen-1 was a depth-conditioned video model, so it would turn. it would take a input video, it would predict. it would it would first convert it into the depth map, and then we would generate, pixels with a latent diffusion model.

Swyx [00:11:01]: Yeah, very effective.

Vibhu [00:11:02]: Yeah. I didn’t realize how distracting the blog post would be. Sorry.

Anastasis [00:11:05]: Yeah, but, one of my favorite examples of on those, on Gen-1 was both, if you go up to mode three or mode two, there was this storyboard use case where people would make

Vibhu [00:11:18]: Ooh

Anastasis [00:11:18]: Would

Vibhu [00:11:20]: You can mess around with the

Anastasis [00:11:20]: Make a city out of books or out of boxes, and then they would shoot a video with their phone and then translate it into a photo-photorealistic output. There was all these ways in which those models were starting to be used for storyboarding and also for really. and then if you go to mode four, like, of taking untextured 3D scenes and then turning them into photorealistic output. So we saw a lot of use cases early on where people that were familiar, were power VFX editors would just take a blender, render, and then they would get translated in with Gen-1 or create a scene in Unity and then take a capture a video of it and then translate into, restylize it. So I still think video to video is powerful. I think we had a recent video-to-video model as well, and it’s one of my favorite ways of using those models is essentially using them to use ground truth video as, like, the initial inspiration and then translate into different styles or different outputs.

Stable Diffusion, Stability AI, and Open Source

Swyx [00:12:23]: But I think we’re gonna go into, like, the rest of Runway and catch people up to speed today. I did wanna cover the, let’s call it the Stable Diffusion controversy, or, what happened with Stability AI, whatever. I think there was a two sides of the story. I think there’s part of that is a normal thing of, like, people, join and leave companies, but what is the, retrospective now that, there’s been some years behind it?

Anastasis [00:12:49]: Yeah, it’s a very, it’s a very long story to go into. I think it would

Swyx [00:12:53]: Which I remember you wrote a really long post about.

Anastasis [00:12:56]: We would probably cover the whole hour to go into it in more detail. But, essentially, there was the latent diffusion paper that came in, I think that was at the end of, 2021. And then Patrick Esser, who was one of the researchers behind, latent diffusion, and he worked at Runway at the time, he built latent diffusion in collaboration with Robin Rumbach and a few other folks back, in the in, CompVis, which was, a lab

Swyx [00:13:26]: Like a research group, yeah.

Anastasis [00:13:27]: And, after releasing the early latent diffusion model, they, essentially they were. the goal was to keep working on versions of the model, scale it up, incorporate new data, incorporate new tasks. And Stable Diffusion was the same model, but trained on more compute, and then with a few more tricks, like a classifier-free guidance paper came at some point, I think in the early 2022. And that

Swyx [00:13:52]: Which, like, was a big prompting improvement.

Anastasis [00:13:55]: Yeah.

Swyx [00:13:55]:?

Anastasis [00:13:56]: That improved results. it was trained on better data, so like, the esthetic subset of LAION, but it was effectively, the same underlying architecture. And there was that big training run, that, happened on Stability’s cluster. Stability financed that run. And looking back at that story, I think it was the work to build and train that model was done. It was a, it was a research project. It was done as part of, like, continuation of the latent diffusion work. It then, I think it the model became very successful, and it, I think there were the. And I think as a result of its success, other companies tried to, figure out the commercialization path for it. But for us, it was very important that we try to, we make sure that we. It was meant to be an open source research project, and so the we decided that we should continue releasing versions of it, since that was the original goal of Stable Diffusion, and that led to releasing Stable Diffusion 1.5. There was maybe a day of, a bit of, miscommunication there, but ultimately that was resolved very quickly within hours. so yeah, there was

Swyx [00:15:12]: Okay

Anastasis [00:15:12]: Not a nice

Swyx [00:15:13]: I just wanted to. you have to

Anastasis [00:15:15]: Yeah.

Swyx [00:15:15]: You’re one of the main players in that journey, and so it’s nice to hear from the source of, like, what happened. Yeah.

Anastasis [00:15:22]: Yeah. I think it’s all, it’s all in the past now

Swyx [00:15:26]: Yeah

Anastasis [00:15:26]: I would say. and, like, both companies, Stability took its own path, Runway took its own path.

Swyx [00:15:32]: Yeah. There’s still. James Cameron is backing the new Stability, whatever they’re doing with the Hollywood studios.

Anastasis [00:15:38]: Right.

Swyx [00:15:38]: I don’t know what they are doing. I think one thing that impresses me, and I’m happy to move on, is that back in the that time, let’s say, like 2021, 2022, there was this community of people that you were involved in that was researching all this stuff, right? And, like, from everyone I talked to who was active then, it seemed like it was fairly obvious that somebody would do the hero training run that would produce Stable Diffusion. So, like, I guess the question is, like, you had the you were you had made investments. You were you had the foresight. Is it accurate to say, like, that is reflective of, like, what people were thinking at the time? Or was it still very much like, “Well, we’ll use it as, like, a post-production tool or something. I don’t know.”? Like, where in the sentiment were we that maybe you can think back to, like, what the community was like back then?

The Early Creative AI Community

Anastasis [00:16:28]: I reminisce and I think very fondly those early years, from like 2018 to 2022, because it was a very small community that, as you said, were very convinced that this was gonna be a big thing. And at the time, anyone who. Because it was such a small circle and, everyone who would, like, be part of that circle and, like, make projects with it would, immediately get, go viral. so like

Swyx [00:16:55]: And you didn’t know who they are, right? They’re just some name on a, GitHub or Hugging Face somewhere.

Anastasis [00:16:59]: Exactly, yeah. So I remember one of the first big viral moments of creative AI was, there was the neural style transfer paper

Swyx [00:17:09]: Huh

Anastasis [00:17:09]: That

Swyx [00:17:10]: Something dreaming?

Anastasis [00:17:11]: I think it was called neural style transfer.

Swyx [00:17:14]: Okay.

Anastasis [00:17:14]: There was also Deep Dream, the puppy slice

Swyx [00:17:16]: Yes

Anastasis [00:17:16]: Which was, also really cool. but, yeah, there was this project that, Jim Kogan, who was an early advisor of Runway and one of those,

Swyx [00:17:25]: Marketing guys

Anastasis [00:17:26]: Big, creative AI, folks, he literally just, like, showed a video of himself taking the New York Subway and going over the Williamsburg Bridge and then stylized it with, I think in the style of Van Gogh or, like, one, painter. And that was. Like, at the time, that was, like, so cool and it went viral and it was completely revelation to people that you could do this with generative models. And that was only, it was less than. It was maybe 10 years ago. So just, like, as an indication of, like, how quickly things have gone.

Vibhu [00:18:02]: It’s pretty crazy. Like, even since then, you’ve got people at every level of the stack. You’ve got devs, creatives, artists, hobbyists. You’ve got everyone using it. And for people that tried stuff early, they’ll remember how hard it was to use regular diffusion, right? Like, nowadays, you can use your favorite ChatGPT image gen or whatever, give a sentence, get a beautiful output. But diffusion was like, the whole ultra HD, 4K, high resolution. Like, prompting these things was very different. anything you learned on the tooling side, like from the offerings you guys have now, so like creatives, devs, you really took the. Research and brought it to everyone to use. anything interesting there to share?

From Gen-2 to Controllable Video Generation

Anastasis [00:18:44]: We had to build the entire model serving infrastructure for video diffusion models. There was nothing else, already, like, because we had Gen-2 was the first text-to-video model, I think, out in the market. So many things that we learn over time. I think the I think the biggest one was, like, we. it was very clear early on that text-to-video was not gonna be the answer. Like, you. Like, people wanted a lot more control than that, and so we invested in, like, control building on top of those models very quickly. how do you use the camera trajectory as control? How do you use an initial input frame as control? So that was a very early learning for us. With text-to-video was, like Gen-2 was an amazing, step function improvement in the quality of video models, but it was used much more in an exploratory way because there was nothing to ground it to. There was no reference that you could bring into it. There was no. You couldn’t really control the camera motion. You couldn’t control the object motion. And so the first year, in 2023, was really all about what are all the interesting ways in which we can condition those models? And it was a lot of just post-training rounds on top of the base model to figure out, like, what, -- how do people wanna control them? And so there was, like, this quick succession of the we it was called Motion Brush, which was you could, like, you could draw arrows and dictate where things should move in the scene.

Vibhu [00:20:09]: That’s so cool.

Anastasis [00:20:09]: There was camera control that was you could just describe, like, how you want the camera to move in the scene. And because we work with filmmakers from the most of the history of Runway, we immediately got this feedback and got this, decided that this was worth investing in. And so control ability became a big theme, I think, very early on as we were building, as we were building those models. Something fun that I haven’t really talked about too much was just how Gen-2 came to be out of Gen-1. So it was a bit strange because we announced Gen-2 two months after Gen-1 and

How Gen-2 Came From a Weekend Hack

Vibhu [00:20:43]: We’re accelerating.

Anastasis [00:20:44]: It was before Gen-1 was even generally available. But Gen-1 was a depth-to-video model, so it would take a depth map and it would convert it into RGB. and we couldn’t get, text or image-to-video to work directly, and that’s why we started from depth to video. but, and we had discussions of like, okay, we need to spend the next six months investing in text-to-video, maybe increasing the compute scale or the model scale, like train a larger model. And I had this weekend project idea, which was, what if I take a model that, starts from text input and converts to depth maps and then use Gen-1 to convert the depth maps Into RGB?

Vibhu [00:21:29]: It would probably work.

Anastasis [00:21:30]: And so Gen-2 was that.

Vibhu [00:21:32]: Oh. The hackathon pipeline.

Swyx [00:21:35]: The weekend hackathon pipeline.

Anastasis [00:21:36]: Yeah.

Vibhu [00:21:37]: But it looks good.

Anastasis [00:21:38]: And it worked pretty well. there were if you, with the knowledge that it has this, like, two-stage pipeline, you can tell in some cases that the structure of the video looks a bit off because you had to generate the depth first before you go into the output video. But it worked and it allowed us to bring this to our, to users very quickly. But it’s now it’s interesting because, like, people are coming back to this almost two-stage approach. Like, if you look at the Reve text-to-image model that came a few months ago, it had this planner model that would generate bounding boxes before it fed that into the diffusion transformer.

Swyx [00:22:19]: Yeah, Ideogram also the same day.

Anastasis [00:22:22]: Yeah.

Swyx [00:22:22]: I remember that was very strange that both of them came out the same day with the same exact innovation.

Anastasis [00:22:26]: It’s a small community, I think.

Swyx [00:22:28]: I’m like, this is like, this is completely coincidental, right?

Anastasis [00:22:32]: People talk. So yeah, there’s, there’s definitely something into this approach. And, now, like every single like, video generation model in production uses a complex prompt completion pipeline under the hood. I think that’s no secret that there is. That

Swyx [00:22:48]: Humans are terrible at prompting.

Prompt Rewriting, Camera Control, and the Seed of World Models

Vibhu [00:22:51]: I think across the board.

Anastasis [00:22:51]: Yes.

Vibhu [00:22:52]: But yeah, I think like the original Sora one blog post even told you that what happens after your input is rewriting your prompt. It’s much more descriptive about what you would want.

Anastasis [00:23:02]: Exactly. I, And there was the DALL-E 3 paper beforehand that, was the first public, description of the fact that synthetic captions and really detailed captions work really well. And then Sora built on that. Yeah, so it was 2023. We were releasing all these updates to Gen-2, like the camera control, Motion Brush. And there was something very interesting about camera control because it was the first time that you felt that instead of, like, you were creating video, you were creating a short video, you were navigating inside the world. And I think camera control was maybe the seed of some of the ideas that we had around world models and really opening up that research direction. We realized, it was this era and this series of, Gen-1 and Gen-2 models really proved to ourselves, yeah, this is the

Swyx [00:23:56]: Cool.

Anastasis [00:23:57]: So this is not the original camera control. This was the updated camera control on top of Gen-3. But yeah, I think it made those models usable to filmmakers, I would say. The so camera control was very popular. And so we realized, there is one way of seeing those models, which is, you’re just as content creation machines, and there is the other way, which is you’re. As you’re predicting video in order to predict video well, you need to simulate the world in an increasing and increasing capacity. And if scaling laws apply on video, just like they apply on language models, then as we scale the compute that we put into those models, then they’re gonna be able to simulate physics, they’re gonna be able to simulate human actions and dynamics increasingly well and predictably well. That was the thesis about around our efforts on world models, and we spin up this research group to just focus on the world models and how do we turn the video generation models that we’re building into something broader and something that would be useful beyond, also content creation as well.

Swyx [00:25:04]: And that was roughly when?

Anastasis [00:25:06]: Yeah, so that was in

Swyx [00:25:06]: Oh

Anastasis [00:25:07]: In late 2023.

Vibhu [00:25:08]: Interesting. like, I think, a lot of people have been saying a lot of video gen model companies have all pivoted to world models these days, but like, 2023, you’re posting it. one

World Models: From Video Generation to Simulation

Swyx [00:25:21]: It’s, it’s debatable whether it’s a pivot.

Vibhu [00:25:23]: Yeah.

Swyx [00:25:23]: Like, arguably

Vibhu [00:25:24]: Yeah

Swyx [00:25:24]: That’s what you always had to do anyway, right?

Anastasis [00:25:26]: It’s in a way an expansion

Vibhu [00:25:28]: Yeah

Anastasis [00:25:28]: Of the applications

Vibhu [00:25:29]: Yeah

Anastasis [00:25:29]: Of the models as they become more capable.

Vibhu [00:25:31]: The early signs, it seems like the original models you guy had, guys had, people would say it’s very not bitter lesson pilled, right? You’re adding, rewriting prompts, you’re having all these one-off things, but that’s just the state of the tech as it was versus the future of as you said, you can scale it up as, we can scale up to world models.

Anastasis [00:25:50]: Yeah. So it just became. And if you looked at the outputs of Gen-2

Vibhu [00:25:56]: Yeah

Anastasis [00:25:56]: It was not. I think it was not obvious to people that this would scale to become a general simulator of the world. Like, you had very limited movement, you had, very low fidelity or low resolution, like obvious mistakes in human anatomy, like all kinds of limitations. But it was just, the idea was that’s just GPT-two, and GPT-two, it can barely generate, like, coherent sentences. Similar, Gen-2 can barely create coherent video, but if you scale it up, you’re gonna. There is no reason why it shouldn’t work in a way. It’s, And I think that was. That’s, that’s always the mindset of Runway is like this extrapolation of, like, if, like, even when we started in 2018 and you looked at the results of the day, you need to look more at the trend of, like, where we were in 2018 versus when we were at the, when the first GAN came out in twenty, four 2014 or twenty, fifteen. And, you started from, like, thirty-two by thirty-two images of faces, and then by the time in 2018, you could generate, street images at the 1K resolution. And it was the same with world models, very early signs of something much bigger.

Swyx [00:27:08]: Yeah. I was gonna say, like, it’s diffusing into focus. Like, if you look at our visible output from year to year, it looks like a diffusion process itself.

Anastasis [00:27:17]: Yeah.

Vibhu [00:27:17]: Especially watching the early, like, old blog posts, you can really see the choppiness, the details.

Anastasis [00:27:24]: Yeah. Like human civilization starting from random noise and then

Vibhu [00:27:27]: Yeah

Anastasis [00:27:27]: Denoising into

Swyx [00:27:28]: Yeah. Just run it a hundred years.

Anastasis [00:27:30]: Civilization.

Swyx [00:27:30]: Yeah.

Vibhu [00:27:31]: That’s how you’re on track, you’re still noising, right?

Swyx [00:27:34]: Yeah. I like the way that you guys phrased it when you, announced it in June, which is, oh, that you had a video essay. “The human mind is no longer the center of AI. Our world is.” Right? Which is, let’s, let’s call it the past five years of LLM-based AI is very much like trying to emulate human preferences and human speech. But now that’s, like, mostly solved. I think that’s, like, some of the context of your essay, which you also wrote around the time. And now it’s like the focus is on modeling the world accurately.

Scaling Laws for Video and Why Predicting Pixels Matters

Anastasis [00:28:03]: Exactly, yeah. So the way we see it is, there is that, initial mission statement of DeepMind, which is, solve intelligence and then use it to solve everything else. But I think it’s starting from everything else, could be valuable of, like, starting from. there is just so much complexity, and detail in the world that in order to. That it’s, it’s hard to learn directly from just human descriptions of the world. Like, we’re assuming that, like, language models learn from everything that humans have written about the world, like our own understanding as of, the twenty twenties. And there is just so much that we don’t know and so much that’s not captured by existing text, about both the low level dynamics of the world, like we’re not describing in detail. if I tell you to describe, like, how do you tie your shoes, that’s a very difficult thing to describe in words, but it’s very obvious thing to demonstrate. And so I think there’s been. And there’s, more of X paradox, like we’re constantly underestimating all the complexity that goes into very, like, things that we do subconsciously as humans, and we don’t even necessarily always have the words to describe them. And so in my mind, the simulating the world and simulating, physics, simulating the dynamics of the world has always been underestimated, compared to, we place too much emphasis on the things that are easy to talk about. but there is just all this complexity and richness of the world that if we just try and train directly on that observational data instead of training on how people describe the world, we would learn something new that we wouldn’t otherwise know.

Swyx [00:29:54]: You think that the present architectural paradigm is fine? You don’t need, like, another layer, like JEPA, like another famous, New York AI leader would say?

Anastasis [00:30:05]: We’re a very pragmatic research lab. If, we have evidence that an approach works better than the approach that we’re taking, then we have no qualms to taking it. We just have seen no indication that video prediction itself doesn’t scale. And even if you look now, not just our work, but the work of others, you’re seeing in robotics some of the most promising work, starts from video prediction models, and then you adapt them to also the action models, for example. so there is very little evidence that you need something else and that your time is better spent on a novel architectural change compared to improving data and improving the, and scaling the current approach. And so, We don’t have any indication that. the, there is that counterargument that I think there was a tweet by Yann LeCun a few days ago that, understanding the dynamics of the world is very different than, generating, cute videos.

Swyx [00:31:05]: And your answer is no, they’re the same thing.

Anastasis [00:31:07]: Yeah, they’re the same thing.

Swyx [00:31:08]: My cat videos are the same as understanding physics.

Anastasis [00:31:11]: Right, because if you wanna generate. video models can cheat and, like, they could you could give, like, successive dif shots of the scene in a way that doesn’t require you to simulate difficult physics. There is like, all these different ways in which you can hide the deficiencies of the model, and it’s important not to be too tricked by the performance of the current video models. It’s easy to, cherry-pick examples and think that video models are further advanced than they are. So there is a lot more work that we need to do to improve those models. But in my mind, very similar to language, and, like, we’ve. you go from barely coherent sentences to something that, could hold a conversation with a human to something that could can operate autonomously for a day and, like, create entire code bases. And the main difference, there is some architecture improvements along the way, but the main thing is scale. And so it’s the same bet for video, and we have no indications that this is saturating. Like, we have benchmarks that we use for measuring the physics of those models, and we see those predictably improve as we scale those models. So there is. If you want to Google up, Physics-IQ, is one of those benchmarks that measures how well does the model perform at solid mechanics or fluid dynamics or optics.

Vibhu [00:32:32]: I’m curious if you’ve seen any emergence, any scaling law around this.

Swyx [00:32:37]: Yeah, he’s saying there is a scaling law, right?

Anastasis [00:32:39]: Exactly.

Vibhu [00:32:40]: Yeah,

Anastasis [00:32:40]: So the way those models, those benchmarks work is you. the researchers have gone and, like, captured, a few videos that are representative of different physical phenomena, and then you can take the first frame and then pass it through an image-to-video model and then generate a rollout that shows what should happen next. So you have, a ball hanging from the ceiling, and then you use that as input, and then you the model predicts how the ball should fall on the ground. and this measures. we have an intuitive understanding of physics. I know, you can imagine what will happen next if I drop this bottle. So it’s measuring that same intuitive physics understanding of those models, and we’ve measured that at different model scales, and we see, and compute scales, and we see that the score on physics IQ predictably improves. There’s other, tricks and techniques that you can make to improve the score even further, but even scale alone helps, in the model learning better physics.

Swyx [00:33:40]: My main sympathy with Yann LeCun is the, Plato’s cave allegory, right? Like, you’re, you’re, like, learning on the output of a thing, not the internal process of a thing, and it’s very noisy. And, if only you could observe the internals of a thing. It’s hard to observe the internals of a human mind, but you can very much observe, or at least we have a whole branch of science and physics that we’re ignoring on how to model Physics and movement and, gravity and, other interactions. and we’re just, like, throwing away all of that and just saying just scale data, which is very much the lesson of unsupervised learning, but it feels wrong. that’s the main idea.

Anastasis [00:34:21]: I think the history of machine learning is, at large, it feels wrong.

Swyx [00:34:25]: Yeah. It’s a bitter lesson, right? Yeah. It’s, it’s, it’s the simple answer to that.

Vibhu [00:34:29]: I guess, how much can you scale? So, like, even on, let’s say, the video generation side, like, there’s one side of video understanding. Video generation, are we still gonna have tools where it’s like, I wanna generate two hours, twenty hours? there’s a infra way to do it in batches and stitch it together, but, like, do we just keep scaling? Do we just continue long generation consistency, all that at scale? And, like, tying it into where we’re at now from we looked at Runway two to four point five

Gen-3, Sora, and Runway’s Scaling Inflection

Anastasis [00:34:58]: Yeah.

Vibhu [00:34:58]: Like, technically, what advancements have we made to today, and then where do you see things still going?

Anastasis [00:35:04]: So part of the answer is definitely scale. and that was. We learned that lesson in a big way for with Gen-3. So Gen-3 was the model we released the year after, like in 2024. That was a few months after Sora was released. so yeah, there’s an interesting story of that came to be as well. Gen-3 for us was, the first time that we really needed to build. we had to learn all the lessons that the language model world learned in two in three years in the span of a few months. one of the biggest changes of Sora was using diffusion transformers instead of convnets. So a lot of the early, latent diffusion models were all, convnets for the diffusion model part. And the diffusion transformer paper came at some point in 2023, and it showed scaling laws for image, diffusion transformers. And we realized at that point that we needed to invest in infrastructure for model parallelism, for really scaling training to larger than, a few billion parameter models. And we spent maybe the, most of the fall of 2023 building out our infrastructure for distributed training. And we had a lot of false starts and a lot of failure in trying to scale, image and video diffusion transformers. And at that point, February 2024, Sora comes out, and the results are

Anastasis [00:36:35]: Very much superior to what Gen-2 could produce. There were a lot of, a lot of chatter on Twitter about Runway. Runway’s done. like, there is no way Runway will catch up. And if you remember, also OpenAI in the early twenty-It felt very, like it’s a

Swyx [00:36:56]: To the moon

Anastasis [00:36:57]: It’s a formidable opponent now, but at that point, it, they were on the top of their game. nobody could even get close to them. There was maybe Gemini was just the first version of Gemini had just released. So when OpenAI came with Sora and it was such a big jump of like quality, it gave me, there was like an existential crisis for a few hours. But that, I think the amazing thing about Runway and like I think the, we’ve been around eight years now, which is almost we’re dinosaur in AI, and we had to like, we had there was a lot of those moments we had to learn, adapt very quickly and build out skill set in the team that we didn’t have. And so, if you ask anyone what is their favorite time at Runway that was there during that time, it was that push in like three months to get to a model better than Sora. and it, we scaled 10x the model scale, the model size and the, compute that we were training on. we figured out model parallelism. We had zero expertise in that. And then we came out with Gen-3 during that summer. So that was a big turning point, I think, for the company where the research org grew very quickly, and we really started pursuing this vision of the general world model, in earnest, I think after Gen-3 was out.

Swyx [00:38:12]: Yeah. that’s the amazing thing about building when you’re building. There’s no stack to. You have to invent everything yourself. You have to be completely full stack. Now I think like there are inference specialists like Fal or whatever that can help with like, model serving, and I think you guys work with them as well. but yeah, like it’s, it. But at the time, it was just. It’s very interesting to think about what you do when Sora comes out and people are questioning whether your company should still exist.

Distillation, Turbo Models, and Real-Time Video

Anastasis [00:38:41]: Yeah. And yeah, there was no, there was no VLM of diffusion models. Like, we had to build the whole model serving infrastructure and make things efficient. And a few months after we released Gen-3, we released the Turbo version, which I think was the first step-distilled model in production.

Swyx [00:38:56]: That was a whole trend that we covered as well. Yeah.

Anastasis [00:38:59]: So that allowed us, to serve those models at the larger scale, ‘cause I think the first version of Gen-3 was quite, expensive to serve.

Swyx [00:39:09]: I think the whole like trend in like consistency models, Lightning and, Turbo and all these things somehow didn’t really stick around. I don’t know if you have any reflections on this. Because at the time, I was like, “Well, everything should start with a distilled model first, and then you can upscale,” right? It. your bigger models just turn into fancy upscalers, but like you should always draft with a smaller model and faster model, right? Because you can get it so quickly, like near real-time.

Anastasis [00:39:39]: Yeah. I would not be so sure to say that didn’t stick around. I think that, it’s, it’s likely to. that there is a lot of step-distilled models that are actively used in production. there is still a gap in quality compared to the, non-distilled model. but in my mind, we’re still. there is a two to three year offset from language models. So the things that, So it’s just a matter of time before there is better distillation techniques. we use. Right now we have a real-time model core character that I think is the largest deployment of real-time video models, that’s a step-distilled model, and it’s actively being used. It’s a very specific use case compared to a general video model. So this is a

Swyx [00:40:27]: Very cool, by the way.

Anastasis [00:40:27]: This is avatars stuff, right?

Swyx [00:40:28]: Consistency, character.

Anastasis [00:40:30]: Yeah. So this is a talking avatar, model. we were able to. we optimized the hell out of it, and it generates at 24 FPS, and it’s a, it’s a step-distilled autoregressive video model. So if we look at our world model direction, a big component of it is starting from the bidirectional diffusion that generates entire video at once and making autoregressive shows. So you generate one frame or a few frames at a time. so there’s a lot that goes into that pipeline of getting to a real-time model. It’s first you need to make it into a causal autoregressive model, and then you just turn it into. You need to do some additional step distillation to get it to be real-time. and I think that part is just starting. I’ll be very surprised if we’re, two years from now, we don’t primarily use real-time models. To me, real-time video generation is just inevitable that, it has much better user experience, it’s much cheaper to serve, and, the quality gap between the base model and the real-time model is only gonna close as we figure out better, distillation techniques. And we made a lot of progress there internally on maintaining the quality of the base model when we distill them.

Swyx [00:41:49]: How much of this is transferable? So is it the same base model? Like if you’re doing diffusion across the whole sequence and you’re converting it to step autoregressive distillation, is this like distillation where you still need to train both, you can use the same base and converter? What’s that process like to go from regular model to something that’s real-time on a technical level?

Anastasis [00:42:11]: So the nice thing about diffusion models is you have, two axes of distillation. So there is the. You can distill to a smaller model, which resembles what you do in LLMs, or you can distill in terms of taking less steps, less diffusion steps. So you could take a model that generates in fifty steps and generate in four steps and get to, You have some performance, degradation, but very often you get comparable outputs. So you can even take the large frontier model and distill it with step distillation and get to a real-time performance, and that’s what we’ve seen. So, depending on the use case, in some cases we might also serve with a smaller model, but in a lot of use cases, we just use the

Swyx [00:42:56]: Step distillation

Anastasis [00:42:56]: The frontier model, and we’re able to make it work in real-time.

Swyx [00:42:59]: I think this might be a good time to cut over to his laptop to show off some of the real-time stuff that you’re doing.

Interface World Models and Neural Software

Anastasis [00:43:06]: This is one of the research updates that we did recently. so we’ve been working and f in getting our general world models to, different applications. one of them that we think is very compelling is using general world models as essentially, an interface, a universal interface to software. This is a version of our world model that’s called an interface world model. and the idea is that it essentially, replaces, the, front end of a software application. It renders the pixels directly of an interface and is trained to predict what happens next as a result of, a click or another interaction you have with the interface. So this is all pixels. it’s there is no HTML, CSS, React that’s powering this interface. This is directly at the output of our real-time, video generation model, and it takes clicks directly as input.

Swyx [00:44:09]: And drags, click and drag.

Anastasis [00:44:12]: Right. So it supports

Swyx [00:44:13]: Ooh.

Anastasis [00:44:14]: Yeah, clicks. It supports drags. it also supports scrolling. and the amazing thing about this is that you can effectively describe in the prompt how you want different elements, like what do you want the behavior of different elements to be. So it’s almost you’re you can turn, an interface from, markup language description of, like, an HTML interface, and instead you can just describe the interface. if I press this button, I expect this to happen. If I press this button, this should happen. And it’s useful, we believe, both for prototyping, for, like, just testing, like, what different interactions would feel like. you can also add audio to it. So it’s a video audio generation model. So you get you essentially can describe both what the visual outcome should be of your click and also what the if there is a sound effect that comes out of it. So we believe that’s gonna be a much more flexible way of building software. Just render. It just, in why generate the code that generates the pixels? Just generate the pixels directly.

Anastasis [00:45:18]: It’s the end-to-end philosophy applying applied to front ends.

Anastasis [00:45:25]: So we think there is a few interesting use case. So you can build creative tools on top of it.

Anastasis [00:45:32]: We think that, for any use case that involves a lot of exploration or, like, educational use case where you wanna learn about a new concept and you want some visualization and like, and open-ended exploration, we think those this is a very powerful, approach. you can imagine new forms of, design, industrial design software that could emerge as a result of those models. And this is all, generated in real-time as well. So, you can build a lot of interesting camera transitions and forms of interaction that are very difficult to build otherwise. And one way in which we evaluate this is what if you try to generate the same interface with Claude by just, prompting Claude, “Here’s an image reference of my interface that I made in Figma or that I created somewhere else. create this particular interaction,” which in this case it’s, drag that object, upwards. and beyond it being slower, it’s also very difficult to capture some interactions by just fully, with just LLMs. So we think that this is likely to be the way that a lot of the future, like, software in the future will be created. and one of the additional benefits is personalization might be a lot easier done with those models. Like, you can essentially try out different prompts based on who is visiting the interface. You can, more easily, prompt engineer the interface to have larger size, text for more accessibility reasons, or you can make this or, like, if you have a particular aesthetic preferences. So we’re very excited about this approach. It’s early days, and I think we’ll need to, make it more cost-effective as well to serve those models ‘cause, running a real-time video model versus just purely rendering HTML, there’s -- the computational needs are much higher. but we do see a lot of potential in this approach to building front-end interfaces.

Swyx [00:47:47]: So we covered this similar thing with Flipbook before with our, Ethan Hara episode with Groq, video. And yeah, I think it’s very engaging visually. I think it’s maybe very good for education, but it’s it does sound expensive. I think there’s an upper bound to how expensive it will be, though, right? Like, the inference cost will go down over time. You’ll figure out ways to optimize it. Effectively, when it pauses, you don’t you’re not receiving human input. You don’t have to generate anything, right? So.

Anastasis [00:48:14]: Yeah, you could also. Like, in this case, you have ambient motion, so there is parts of the screen that might. if you’re let’s say you wanna, visit Paris and then you get this interface that allows you to explore.

Swyx [00:48:29]: People walking. Yeah.

Anastasis [00:48:29]: You have people walking or, like, things happening. But, it’s, it’s a no Yeah, it makes it more expensive because you need to run the model all the time. Maybe you have some looping mechanism so you don’t need to do that. But all those things, I think, is stuff we’ll need to figure out.

Toward a Fully Neural Operating System

Swyx [00:48:44]: Yeah.

Anastasis [00:48:44]: I think our first consideration is let’s make this clearly find some use cases where it’s clearly a much more compelling interaction compared to traditional interfaces. And then it’s a matter of time before it becomes more cost-effective to serve.

Swyx [00:48:58]: Yeah. When it comes to the people walking, I think the approach that makes the most sense to me is Nick.

Anastasis [00:49:04]: Nick.

Swyx [00:49:04]: Oh, God. I keep messing up their name. With Chris Manning and Fanny Yan. I don’t know if you’ve come across them, where they. Mapped to some game engine. I think it’s Unity or something, or Godot. And they you can script some NPC behavior behind that and train on that. Whereas here, you can really imagine whatever you want. Like, that is a UI, right? Like, and it feels, like, more tractable, I guess, to, create a world model of software that is interactable because we have many of examples of that, and you can, do your fancy RL environment stuff on that than it is scaling up to embodied and real-world physical use cases. But this is a nice first step.

Vibhu [00:49:43]: Or, there’s the opposite of you have, like, one B models, three 50 million parameter language models. It just gets so small that they’re just predicting, like, fishes moving.

Swyx [00:49:53]: Small models are now 120 B, so.

Vibhu [00:49:57]: Ultra mini on device.

Vibhu [00:49:58]: But, no, I think it, like, it puts it into perspective, at least the car one for me, like, the applications, right? The amount of work to do that, sure, you only make one model year car per year, but applying this, it’s also a cost-saving to have to manually make all this, right? So it opens up a lot of possibilities, too. I’m curious if you extend this out two, three years, so where do you see things going even further?

Anastasis [00:50:25]: Effectively, the end game of something like interface world models is you have, a fully neural operating system. So I think, Andrej Karpathy has written about that quite a while back. But it’s, You, I think to me it’s, it’s a bit, it’s a bit odd that, we have, for example, with an interaction with an LLM of today, you have this LLM that can talk to you about anything. It can You can take the conversation in any direction. You can It’s very general, so it can solve all those different tasks, but you interact with it through a very rigid interface. And so to me, it’s just a matter of time before the interface itself becomes learnable and becomes, part of the whole loop of, like, you’re not just delivering. You’re delivering an application end-to-end, and that means you’re delivering the language model, but you’re also delivering the render and the pixels and that’s also a learnable component. And the concept of applications might not necessarily. I think we’ll need to figure out new abstractions for software. the concept of application comes from this idea that you need, separate code bases to describe, to, for, to power each individual, tool and each individual application. But you might think of something a lot more unified if you’re. if you have, a video model that’s generating the interface as you go. so it can take context from an LLM and allow you to combine different functionalities that traditionally would live in different applications. So it’s a, it’s a way to solve, software end-to-end, effectively. We also see this as a powerful way to train computer use agents as well. so this is, one way to see this as. And in general, with world models, there is those two directions. One is world models for humans and world models for

Swyx [00:52:24]: Agents

Anastasis [00:52:24]: To train agents.

Swyx [00:52:25]: Yeah.

Anastasis [00:52:25]: And so for every new work of, world models that we do, we have this both uses become possible. So this is a powerful synthetic data generator for training computer use models. It could become, a live, RL environment that you could use to do online RL with a computer use agent, and you can get wide diversity of different interactions, kinds of interfaces, just generated on the fly that, to improve the how robust the, your agent, becomes. So that’s the same also with the world models that we’re working on for a robotics use case as well.

Long Context, Error Accumulation, and Autoregressive Video

Swyx [00:53:02]: Is there a research breakthrough that you’re Waiting for that would unlock the next set of use cases that you really wanna pursue?

Anastasis [00:53:10]: Long context is a very important one, so being able to maintain consistency for long periods of time, and that depends on the use case. So for our characters model, for example, or for the interface world model, it’s easier to maintain long sessions of interaction. If you go into more open-ended worlds that you navigate and you take arbitrary actions in, we, like, there is more the context at which you can and duration which you can generate becomes limited much more quickly.

Swyx [00:53:40]: Yeah.

Anastasis [00:53:40]: So we see more degradation and error accumulation happening. so the biggest challenge with autoregressive models is error accumulation, is you’re feeding generative frames back into the model to generate the next The next frames. And if there is any small errors, they accumulate over time. That’s not a new problem. It’s a problem that LLMs also have, and we’ve seen the ability to generate now really long outputs. So it’s a solved problem, but it’s definitely still a challenge.

Swyx [00:54:08]: Yeah. And what is the state of the art? so for Grok, it would be like 10 to 20 seconds of context going in there for video.

Anastasis [00:54:16]: With our characters models, we’re able to generate up to 30 minutes of video autoregressively.

Swyx [00:54:21]: Yeah. But that’s just for the avatars.

Anastasis [00:54:24]: Yeah. So if we look at, GWM Worlds, which is more our open-ended world exploration model, it’s, it’s on the order of a few minutes, which is Yeah, so

Swyx [00:54:35]: Probably enough for people because you have to cut to the next scene anyway, right?

Anastasis [00:54:40]: Yeah, it’s not, it’s not the ideal game experience if you have to restart every few minutes. So I think. But, I think it’s. Yeah, for certain kinds of game experiences, you can work around it. ideally, you are able to just generate forever, and it doesn’t, it doesn’t degrade. And I think that’s a matter of time before we get there.

Swyx [00:54:59]: Yeah. Genie has, like, one, max one minute?

Anastasis [00:55:01]: Right. Yeah.

Vibhu [00:55:02]: This was your. You did a study on robotics. I think I also have just your Runway Robotics page, though. Is this better?

GWM Robotics and Sim-to-Real Evaluation

Anastasis [00:55:11]: So last year we released Gen-4.5, so that was our latest base model. We’ve been As I mentioned, we’ve been doing all this work in world models, and which essentially a lot of our approach to world models is how do you take a bidirectional diffusion model and make it autoregressive and make it accept actions? So instead of being a video you watch, it becomes a simulation that you step in, and you can, control it every step of the way. You can explore counterfactuals, like what happens if I take this action versus if I take this action. And GWM-1 was the it’s the world model that we built on top of Gen-4.5. So we did all this autoregressive and like, distillation, auto-regressive and then step distillation on top of Gen-4.5. And one of the biggest use case that we saw for GWM-1 was in robotics. One thing we like to say is we as we scaled video models, we accidentally, created one a state-of-the-art model for robotics, by just scaling video models. So we realized at some point, mid last year that robotics labs that are coming up to us and asking to use video models for synthetic data, asking us to post-train our video models to work really well for robotics, so that they can use that to generate variations. That was the first use case that we saw. And then increasingly became clear that the models will be useful beyond just creating synthetic data to train robotic policies. They would also be very useful as simulators. So that means that you can use, a video model online to test how your robotic action model performs. So you can take an action role and then get the outcome of the action inside the world model and then continue that loop like this closed loop simulation. And you can use that to evaluate how well your robotics model works. and the biggest thing that I think you need to solve if you want to build a simulator is establishing real-world correlation that if you take an action inside the world model, if you take the same action in the real-world, you get a similar outcome. So that was the goal of some work that we did earlier this year. So if you go to the first link. So that was, essentially wanted to establish that, real to sim correlation for our world model, so that if you do a series of actions inside the world model and if you do the same actions in the real-world, you get similar outcomes. And we took our GWM-1 model and we used some benchmark data that there is this Roborina, benchmark that’s very commonly used to evaluate how well do different action models perform. And we use the same scenarios and settings and embodiments inside our world model, and we measure the correlation of how well did the action model perform inside the world model versus in the real-world. And we saw that we could get very good correlation between our world model and reality. And that means that if you want to evaluate how well your robotic policies perform, you can scale that much faster inside simulation instead of having to do that with actual physical hardware. And so that was a first indication that our models could be, quite useful in robotics. And we saw as we were working with robotics labs that became like the first use case where they could use video models in a way that feed into their training pipeline.

Vibhu [00:58:40]: Can I ask what

Anastasis [00:58:41]: Yeah

Vibhu [00:58:41]: The difference was from four point five to solving that? So the sim to real gap has always been the issue, right? You train a robotics model on video data, it doesn’t generalize to real-world, and the simulation had an issue. So seems like you solved it, but how?

Anastasis [00:58:56]: Yeah. So a big problem with simulators is, if you’re trying to simulate rigid objects, like it works quite well if you can describe the physics of objects very accurately, then you’re able to use, Isaac Sim or MuJoCo or one of the traditional simulators. But for more complex interactions with cloth, for example, or, like slippery surfaces, with the all the complexity that you want to be able to solve with the manipulation, with an action model that solves manipulation tasks, it’s very difficult and so time-consuming to build, for each of those environments and each of those tasks, build the simulated version of that, the digital twin of that environment. Whereas with a world model, you just need to provide the first frame and then you just can roll out the policy inside the first frame. So whereas, we compare it to methods that required like 3D scanning an environment and then 3D scanning each individual object before you can now, you can bring that to simulation. whereas with a world model, you just take a picture of the environment and then you’re able to test how your policy performs. Our general thesis on robotics is, there is companies that are leveraging a lot of teleoperation data to train robotics action models. There is now companies that are using, humie data, which is, essentially human, egocentric video where humans use robotic creepers to perform different manipulation tasks. And then there is companies that are focusing on egocentric data, which is, you strap a GoPro on someone’s head and then you capture them performing a task. We think that, and all those are great source of data for training robotics models, but the most plentiful source of video data is third-person video data. It’s And if How do we as humans learn how to perform different tasks? A lot of it is by observing others perform those tasks. We don’t learn from first person. We do some trial and error and like, to learn different things, but. Ultimately, a lot of what we learn how to do in the world, we learn by watching other people do it. And that’s how when you’re pre-training a video model, you’re essentially doing that. It’s a lot of third-person video footage of people performing different tasks in the world, people doing sports, people doing household tasks. And our main thesis is that video pre-training, once you do that, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data. So you require way less teleoperation data, which is very difficult to scale. and even if you look at egocentric data, which is a bit more easy to scale compared to teleoperation data, which requires actual hardware,

Why Third-Person Video Is a Powerful Robotics Pretraining Source

Anastasis [01:01:55]: It’s still three hours of magnitude less of that exists in the world compared to third-person video data out there. And so our thesis is and generally, like the most plentiful source of data will ultimately wins. Third-person video data pre-training is the right starting point for models that, you want them to generalize and be able to deal with new environments, new tasks, things that you haven’t seen during training. That’s the motivation for why we think our models are especially useful in robotics, settings, and we’ve seen that to be the case, as well.

Swyx [01:02:32]: You said pre-training. So maybe it’s like third-person pre-training, first-person SFT? Is there like a curriculum that you can introduce?

Anastasis [01:02:41]: Exactly. So if we look at GWM Worlds, so GW so GWM Robotics. So digitally in robotics, it starts from Gen-4.5.

Vibhu [01:02:49]: It’s the same video diffusion backbone, right?

Post-Training World Models for Robotics Embodiments

Anastasis [01:02:53]: Exactly, yeah. So you start from the base video model, the one you’re using to generate, cats and dogs and other interesting stuff, and then you, fine-tune on a very small number of hours of robotic data. So it’s something on the order of hundreds of hours compared to if you were to pre-train a robotics model. The current pre-trainings go up to, a hundred thousand or like millions of hours of data. And you’re able to get quite good performance, quickly, because the model leverages all the things that it has learned about the world, physics and human dynamics and the tasks that people care about from pre-training. And ultimately, you want those models to generalize. You don’t want to just be able to perform the tasks that it has been doing training. And the diversity of actions and environments that you have with a pre-training video dataset is much larger than, what you can realistically capture manually.

Vibhu [01:03:54]: How is the scale looking like for the post-training? Like, do you still wanna do, is it like roughly ninety percent of the compute in regular video diffusion model and then scale up a lot, or do it like we want different robotic models for different tasks, or just the one base really good world model can also apply to robotics?

Anastasis [01:04:14]: So currently, we are post-training our models for specific, embodiments that we for particular partners. So if they have a particular single-arm robot or a bimanual robot or a humanoid robot, we would post-train our GWM robotics model on their particular dataset. Over time, we see the different variants of GWM unifying. Like, I would expect, if a year from now or two years from now, you have a single world model that can simulate manipulation tasks, it can simulate navigation, which is a lot of the gaming world models are navigational world models. You’re moving around the space, and it will also simulate human behavior. So that’s the character models. So instead of having three different models, you have a single model that’s able to. ideally, you’re able to simulate what it’s like to be in the world. You’re moving around an environment. You’re maybe performing different tasks. you’re talking to other people. And that happens with, the same, a single real-time video model that’s generating that.

Vibhu [01:05:17]: Do you think you can solve self-driving? So if you are learning to drive a car in a simulator, you have a world model. Your robot is car can manipulate so many axes. How far off are you from something like that?

World Action Models, Self-Driving, and Learned Policies

Anastasis [01:05:31]: So world models

Vibhu [01:05:32]: Or a really good ADAS system?

Anastasis [01:05:33]: World models are definitely being applied to, self-driving, research right now, mainly for evaluation use cases, but our focus has been more on robotic manipulation. We’ve done some work on AV, world models as well. but yeah, we do think that world models are and video models are the best starting point for both simulators and also policy and the action models. So that’s, that’s the other side to this, is that once you have a great world model, then you can just add an action head, and it can predict actions as well. One way to think about it is if you take the starting frame of a scene with a robotic arm and you ask, you prompt the model, generate the arm picking up an object, it would And if it generates an accurate enough video, then it should also be able to generate the exact poses, in 3D that the arm should take to perform the same action. So this is the direction that’s now the popular term for it is world action models, which is you’re starting from a video model, and then you’re adding an action head to predict the actions, and it becomes a policy, essentially.

Swyx [01:06:43]: One thing I’m also impressed by is how much data you need to train these kinds of models. You probably can’t say exactly how much, but like, the original, diffusion models, and from what I know, even of the open source Chinese models, it’s not that much data. Isn’t it surprising?

Anastasis [01:07:02]: What do you define as much data?

Swyx [01:07:05]: Yeah, and it just comes, goes in. Is the token count still relevant?

Anastasis [01:07:09]: So it’s a bit more complicated and,

Swyx [01:07:11]: What is just gigabytes, right?

Anastasis [01:07:13]: Yeah, hours of video, right?

Swyx [01:07:15]: Yeah. Yeah. I feel like something that’s interesting is it seems like the, let’s call it tokens to param counts in language models has really, maybe they’re three years ahead or whatever, seems to be a lot higher than, video models still, even though technically video has more information, per bit. I don’t know if it seems intuitive or maybe there’s just a lot of, like the variability between a pixel to the next pixel is not that high. So, like, maybe there’s just a lot of information that is repeated.

Scaling Video Data and the Lucid Dream Test

Anastasis [01:07:47]: My answer would be it’s still very early. Like, the training video models will scale way further than it

Swyx [01:07:55]: Yeah

Anastasis [01:07:55]: Currently is, and you’ll have capabilities that go much further than the current models can do. So one thought experiment that, I like to use, it’s, it’s almost like the Turing test of video models or like the Turing test of world models, go, I call it the lucid dream test. It’s you have a

Swyx [01:08:14]: You mean the actual person lucid dream?

Anastasis [01:08:17]: It comes from this idea

Swyx [01:08:18]: Lucid rains, right?

Vibhu [01:08:19]: Lucid dreams is telling you’re dreaming while you’re

Swyx [01:08:22]: Yeah.

Anastasis [01:08:23]: Yeah, exactly. So lucid dreaming is when you realize you’re

Swyx [01:08:25]: In a dream

Anastasis [01:08:26]: Inside a dream, and then you

Vibhu [01:08:28]: Play around

Anastasis [01:08:28]: Be able to control what happens in

Swyx [01:08:30]: No, there’s also an inference guy called Lucid Rains. Yeah. Or quantization

Anastasis [01:08:33]: Very prolific, person. Yeah. So let’s say you have a VR headset and you’re in a room with and you’re wearing a VR headset, and that VR headset, most of today’s VR headsets have a pass-through mode, so you can see directly what’s in front of you in the world, or you can render something inside the VR headset. And there’s gonna be a point where those interactive real-time video models become good enough where you wear the headset and you’re in the same room and you’re walking around and you’re kinda and you’re interacting with objects. You’re able to move freely in that room and do, and interact with any object. And at the end, someone asks you, “Did you were you using pass-through mode, or were you -- or was this, rendered or generated, footage?” And if you cannot tell for sure if that was what you were seeing as you were interacting with and moving around the world was generated or it was, pass-through mode and was just what was happening in front of you, that’s an indication that the models have become good enough. And we’re not, we’re not close to that yet. And a lot of it is just this idea of really simulating dynamics and counterfactuals well. Like, if you ask a video model to generate a person scoring a goal versus a person failing to score a goal, it would do a better job at scoring the goal because there is a bias from the training distribution. There is a lot more videos of the person succeeding at scoring the goal. But if you have an interactive model, you want it to be able to generate counterfactuals. Like, if I take this action versus this action, you want it to generate equally realistic outcomes. so that’s, I think, the big gap between video models and world models is that idea of the counterfactual generation. And if you want a great model for robotics, you wanna simulate failure very well, because whether you’re using it for evaluation or you’re using it as a in an online RL loop in the future, you wanna be able to have the model try and fail to do things and improve. and so in order to do that, you need to be able to simulate things failing.

Swyx [01:10:42]: This is the only domain where you have too many successful examples and not enough bad examples. Should be easy to generate failure.

Vibhu [01:10:51]: Oddly enough, I think, like, early image video models weren’t good at being human realistic, right? Like, you see aa lot of the high-res 4K, like, professional photography, but not just everyday life, like normal picture, right? Everything looks like it’s professionally generated, like professional pictures, but not just like normal, like, messy cables on a desk.

Swyx [01:11:13]: Okay, so there’s, there’s this stuff. one thing we also covered that you guys have, video agents that you launched. I guess, how does the traditional, let’s call it frontier, like, autoregressive LLMs, like, feed in, to all this? They’re driving ro your robotics models, or are they driving others, your video agents, production, anything where you see the overlap of autoregressive and diffusion, let’s call it?

Counterfactuals, Failure Data, and World Model Evaluation

Anastasis [01:11:41]: Yeah, so harnesses are really important across all those different use cases. So we have this video agent, which is essentially an LLM that is very effective at tool use of different, image models, video models, and helps you through creating a project end-to-end. So, very often in, like, a traditional advertising flow, you have a brief, you start from it, and then you generate some a storyboard, and then you generate the video. A video agent and, or runway agent helps you through that whole process, and it helps you also analyze performance data. For example, how well did this ad perform versus this ad, and then generate me more of the based on those learnings, figure out what to generate. We think that the harness is a very important piece of the pipeline. as I mentioned, all the video production, all the production video models use some prompt completion that happens, and we expect, that to become more and more complex and more, you generate longer and more detailed descriptions before you use the diffusion transformer. I do think eventually, there’s increasingly this unification into omni models where you have the you’re training the models end-to-end to both do autoregressive text prediction and also, diffusion as well. So you’re predicting the next token, of like you’re, you’re maybe using some reasoning and planning of the scene, and then you’re passing it into the diffusion head that’s generating the pixels.

Video Agents, Harnesses, and Omni Models

Swyx [01:13:10]: Yeah. I think currently maybe only Gemini and Qwen do it. I-I’m not sure which of the Chinese models are omni, but yeah, it’s, it’s not, it’s not a very well, popularized modality, I guess.

Vibhu [01:13:25]: It’s an interesting use case when you think about it, right? Because not only do you have to end at like language model reason, diffusion had generate, you don’t have to output there. You can go back in to feed that output to the same model, reason again on improvements, and it can do a lot of loops just in its own. I guess the question is like, do we need that or can we just do agent scaffold, like do it outside the model? Is there a big benefit to doing it in?

Anastasis [01:13:54]: I think there’s generally the trend of something is first done by a harness and then it becomes part of the model, right? So you had the chain of thought prompting where you had to do this super detailed system prompts to

Swyx [01:14:07]: Yeah, step by step

Anastasis [01:14:08]: Get the output. And now the model generates the reasoning trace by itself before it gives you an answer. And in the, in video models similarly, a lot of the video models of the early days were single-shot video models, and you had to use some orchestrator to turn, generate multiple shots in parallel, and then turn it into an actual video.

Swyx [01:14:28]: Or in ComfyUI, just all over the, all these nodes.

Anastasis [01:14:31]: Yeah, like a spaghetti workflow. and now you have multi-shot video generation where you have the you directly generate multiple shots. And there is a benefit to that because then the video model learns some. to generate a single shot well, you need to figure out a lot of stuff about the world. to generate multi-shot video well, you also need to get some, like, video editing instincts. Like, you need to figure out what is the right pacing of shots. And also, LLMs are not that good at it. Like, they’re not that great video editors. If you ask a LLM to take some videos and then auto-create a edited video out of that, it would feel uncanny. So I don’t think LLMs are that good yet at being video editors. And I think there’s benefit to learning that end-to-end. so I would expect, the training generally is the things that, you need the harness for eventually get injected into the model itself, and you learn that end-to-end.

From Harnesses to End-to-End Learned Video Editing

Swyx [01:15:33]: Do you find that you need to hire engineers who can. or researchers who are also artists to infuse that taste, or do you have artists in residence to distill them?

Anastasis [01:15:44]: We have a large creative team that’s very actively involved in the, in training those models, like on the, in every part of the way. And like, how do you caption video as well so that you capture the stuff that you need for, like, the cinematography, the aesthetics, the camera direction in as detailed ways as possible so that you’re able at inference time to elicit that through the model? we have our creative team also does a lot of evaluation of like, what constitutes a usable video out of those models. And so they’re very involved through every part of the process. And I think that’s one of the special things of Runway is just that mix between like creatives and researchers sitting by, side by side and working together to build the next generation of our models. I think that’s been a really important piece to, how we’ve operated as a company.

Swyx [01:16:37]: Yeah. In some senses, though, you can only do this in New York.

Vibhu [01:16:40]: It’s

Swyx [01:16:40]: Maybe, you have other offices, but like, I try to find some poetic, significance in the fact that you are a big New York company.

Anastasis [01:16:49]: As there’s a few parts to being New York. there is that intersection of all those different industries and, like, media, advertising, like

Swyx [01:16:57]: Yeah, this is very advertising.

Anastasis [01:16:59]: The, like the art scene is New York. Not to say anything bad about San Francisco, but, it’s. There is more going on. There is that component, and there’s also, I think we benefit from being outsiders and thinking of things a bit differently, like not being in the same, like, hive mind of,

Swyx [01:17:19]: BВС

Anastasis [01:17:19]: ASI, of Bay Area and, like, taking. and also taking our time to get where we are today. Like, building the, growing the team intentionally and bringing people who are, yeah, both on the creative side and also on the engineering research side. There’s huge talent pool of amazing people in New York, so that hasn’t really been a problem.

Creative Taste, Artist Feedback, and Runway’s New York Advantage

Swyx [01:17:41]: Congrats on everything. what are you hiring for? what should people look forward to, for the future of Runway?

Anastasis [01:17:49]: We’re hiring across the board. I think this is probably the most open roles we’ve ever had in the history of Runway. we’re growing our research team quite significantly. So if you’re, if you’re excited about video models, if you’re excited about world models, if you’re excited especially about robotics, the robotics team, we’re hiring roles in the robotics across, software, hardware, and research. so definitely reach out.

Swyx [01:18:13]: And, a lot of people don’t have direct robotics background, but what should they have, if they want to be useful in robotics?

Anastasis [01:18:21]: So ideally, some experience with learned policies, would be

Swyx [01:18:26]: Just RLs

Anastasis [01:18:27]: Good for robotics. but we tend to hire generalists as a philosophy and, like, people who learn really quickly. but some experience in the, in domain expertise in robotics is something that we’re, we’re definitely looking for the next months. and then we’re scaling the go-to-market team significantly. There is, a wide, like, very active enterprise adoption happening around video models at the moment, and, we’re really trying to, respond to all the demand.

Swyx [01:19:00]: Yeah. Great. You wanna talk about the, open source robotics stuff?

Vibhu [01:19:04]: Sure. It was just random notes we had.

Vibhu [01:19:07]: NVIDIA launched Cosmo. I guess it’s interesting. So, you’re a founding member AI labs to build open source world models in physical AI. - Anything else to talk on here is open research?

Anastasis [01:19:20]: The biggest thing is that, as I mentioned, while models are still, early, like there is still so much that we you can scale and those models further, so much more advancements and things that we can figure out and how to improve those models further. And I think this is, it’s important that some of this research happens in the open and figuring out what is some incentives for different companies to come together to bring some of that research into the open and open source. And so Cosmos Coalition was a initiative that we co-founded with NVIDIA to bring some of that research as open source. And that could mean open weight model releases. It could mean benchmarks that measure physics and things that people care about when building world models. It could mean infrastructure. So really, how do we grow the ecosystem of world models and make that something that also it’s easier for a developer, a researcher that’s just starting out that is excited about world models to contribute to the field.

Hiring, Robotics, and Enterprise Adoption

Swyx [01:20:19]: I think it’s a there’s some amount of like, is this also our response against the Chinese world models that are being released, or is there not part of the consideration?

Anastasis [01:20:29]: I do think it’s, it’s important for NVIDIA models, if you look at the leaderboards of video models, I would say right now the majority of models at the top ten, top twenty are Chinese models. There is, only a handful of companies that are made it to the leaderboard from like the US or the West.

Swyx [01:20:50]: Yeah. We’re doing better with images, but with video we’re very behind, right?

Anastasis [01:20:53]: And so I think it’s definitely important that we invest more broadly as a community to make sure that we can those models can we have competitive models

Swyx [01:21:02]: Yeah

Anastasis [01:21:02]: Out there.

Swyx [01:21:03]: But like what’s to stop us from just distilling from them?

Anastasis [01:21:06]: I don’t know if that’s the best long-term

Swyx [01:21:08]: Not gonna mention that they won’t

Anastasis [01:21:09]: That you’re bounded by the performance that you can. It’s, it’s almost a bit of a pessimistic

Cosmos Coalition and Open World Model Research

Swyx [01:21:14]: Like

Anastasis [01:21:14]: View that you can get better. you can

Swyx [01:21:17]: It’s free data. it’s, you might as well. Like if they’re, they’re doing it for like, for the text language side, they might as well do it for the video side the other way.

Anastasis [01:21:25]: Yeah, I do think we’re, we’re quite capable of training great models

Swyx [01:21:29]: Okay

Anastasis [01:21:30]: Without distillation at the moment. Yeah.

Swyx [01:21:32]: Yeah.

Vibhu [01:21:32]: So anything you have to say on benchmarks and evals? Like, I feel like what I’m hearing is a lot of people really like arenas for video and image models, customers and whatnot as well. They only want the best on the leaderboard, and they refer to arenas a lot more than language models seem to do. But any notes on benchmarks, what’s lacking? How does the average person compare while these both look really hyper-realistic? More than that, outside of we did talk about like robotic simulation, the physics and all that, but anything to say?

Anastasis [01:22:05]: I think it’s the opposite in some ways. I think people, generally creatives and artists and marketers, other like people that are using our platforms, I think rely less on, arena scores. And it’s, it’s just so easy to, generate with a bunch of different models and then compare the results visually. Like one nice thing about image and video models is you can immediately tell with your eyes like what feels good from an aesthetic standpoint. Like any artifacts, any issues with the physics of those models, you can immediately tell. and so that’s it’s easier, I would say, to evaluate, as a human. there is also those models than it is in language models where you have those very complex math and coding and, tests where it becomes a lot more harder, I think, for humans to evaluate and can discriminate between the performance of models at a time. So I think in practice, people just test out the same prompt with a bunch of different models and see what the results look like. And right now in Runway, you can use our models and you can use third-party models as well. So it’s, it’s very easy to do that.

Benchmarks, Arenas, and How Creatives Evaluate Models

Swyx [01:23:12]: Amazing. We’re gonna end with the AI Runway AI Summit. The last societal issue, I guess, I don’t know if this is a thing, is the, you are at the tension between artists and creatives and AI. A lot of people in that community hate AI. the people that are in the Runway community don’t mind using tools. it’s just another brush. But, how have you seen the sentiment change?

Anastasis [01:23:37]: Our perspective, yes, it’s just another branch, brush. It’s just another camera. It’s, it’s the latest of a long generation of tools.

Swyx [01:23:46]: Technology in art.

Anastasis [01:23:47]: Technology.

Swyx [01:23:47]: Yeah.

Anastasis [01:23:47]: And art and technology have evolved together. I think there’s been a pretty significant shift over the past few months, and it came. some of it you can see with a lot of public figures speaking out in favor of AI and being, like in Cannes, you saw a few directors speaking in favor of AI. We had Ron Howard in our film festival. There is, Mark Scorsese also adopting AI models. So you have more of those stories coming out every day of like a well-known figure, speaking in favor of AI. And it’s just a matter of, in my mind, it’s those models are becoming more and more demystified. I would say I have also a bit of a hot take that one of the things that made the initial response to those models maybe a bit more heated than it needed to be was this idea of text to video of, you have a single text description and you get back a -full video.

Artists, AI, and the Evolution of Creative Workflows

Anastasis [01:24:48]: Yeah, there was a misconception. you can generate it to our feature-length film, but the models of today now take a lot of references. They take they are very controllable. And I think when people see a tool that allows, affords many degrees of freedom and control, they respond to it differently. And it matters less that it’s a generative model than the fact that you can steer it to the direction that you want. and so. I think when people look at, complex workflows on top of those models, when they look at, all the ways in which you can steer them and you can provide now with some of the latest models up to fifty references, like the conversation becomes a bit different because it feels much more like a

Swyx [01:25:34]: Storyboard

Anastasis [01:25:35]: A tool

Swyx [01:25:35]: Yeah

Anastasis [01:25:35]: Versus, like, something that a magical entity that figures out, like, the, your entire film for you.

Vibhu [01:25:43]: Any notes on, like, workflows changing for people in the field? Like, I think engineering at least has had a lot of people where they’re like expectations have changed. I’m, ten X, a hundred X more productive, and you can get a lot more done. same thing as, you’re making dev tools for creatives. any notes there? Like, there’s some people that don’t wanna adopt, some that do. Like, anything?

Anastasis [01:26:08]: Yeah. So I think, in terms of, like, what people care about, I see that we have gone through a few stages. So we started from a stage where the main thing that people were looking for was quality. Like, as, we scale those models, the quality improved dramatically. That’s something that people still care about, but it’s, it’s now in addition to controllability, like being able to steer those models with references, with, different kinds of inputs, with storyboards. And now my sense is increasingly people are gonna care about latency more and more. As those models become better, the ability to iterate very quickly becomes more important. And, like, if you can, with a single prompt generate ten different, outputs, like, almost instantly, you can explore way faster than before. And you get some of the magic that characterized the creative tools of the past, like Photoshop was instant. and we lost some of that with generative models. You’re waiting for two minutes to get back a video, and I think we’re gonna bring, some of that back now with the

Vibhu [01:27:08]: Real-time

Anastasis [01:27:08]: Real-time models.

Vibhu [01:27:09]: Yeah. Exciting. And

Latency, Real-Time Generation, and the Future of Creative Tools

Swyx [01:27:11]: Exciting. the last thing we’ll plug is this one, Runway

Vibhu [01:27:14]: Summit

Swyx [01:27:14]: Summit. You’re finally doing this in SF?

Anastasis [01:27:18]: Yeah. So, we’re very excited about this. So this is, in late September thirtieth, we’re doing a summit on, primarily focused on physically high and real-time video generation. We have panelists from NVIDIA, Physical Intelligence, Botco, DeepMind. Yeah, it’s gonna be, I think, a very interesting series of conversations. We try to make the panels really technical and, elicit actual substantive discussion and hopefully some interesting disagreements and interesting debates on things. And, yeah, the there’s tickets available. Hope people can join.

Swyx [01:27:56]: Since you mentioned it, what disagreements and debates should people think about, or do you expect?

Anastasis [01:28:04]: So it’s things like, there is, one debate right now in the robotics world is, VLA’s versus world action models.

Runway AI Summit and the Big World Model Debates

Swyx [01:28:11]: Okay.

Anastasis [01:28:11]: So there is labs that are really betting on one of those two directions. there is like what is the best source of data to train robotics models?

Swyx [01:28:21]: There’s just the third-party, first-party that we talked about.

Anastasis [01:28:24]: Yeah. There is, the people who really believe in further scaling teleop data versus leveraging more large-scale video data. So that, those are some of the. And then there is, the world models debates of predict pixels directly versus something like JEPA versus a more 3D-based, 3D-based approach. so I think we’re at a nice time in world models because there is still that active debate happening on, like, what is the best long-term direction. I feel very strongly that it’s video predict pixels directly and scaling video generation models is the right approach. But it’s, I think there is a lot of interesting, debate happening, by researchers on, like, what is the best path to take.

Swyx [01:29:09]: It’s interesting that it’s all on, like, let’s call it the policy layer and the data model layer. Is the physical side is completely solved? Like, all the sensors, all the actuators, all these things are. We have everything that we need?

Anastasis [01:29:23]: I don’t think that’s, solved either.

Anastasis [01:29:25]: It’s definitely,

Vibhu [01:29:27]: Different problems.

Swyx [01:29:28]: It’s, it’s like

Anastasis [01:29:29]: Yeah

Swyx [01:29:29]: I wanna dream about all these things, and then I get, I buy a robot or I buy, I try to assemble my own, and I can’t even get the motors to, like, work right. Right? Like, and it’s you’re dealing with very sensitive, equipment that has, voltage and power and, like, heat and all these things which, you, abstracted away. We’re sitting here, we’re talking about software and talking about models, but, like, really you have to deal with those kinds of things too.

Anastasis [01:29:56]: Yeah. And, I think I’m, I’m, I’m generally also not opposed to incorporating other modalities into our models like we’ve seen.

Multimodality, ImageBind, and the Maximalist World Model

Swyx [01:30:04]: Yes.

Anastasis [01:30:05]: The simplest case is they can generate video and audio at the same time. So they can generate RGB, and they can also generate, they can generate sound and audio. But my. I’ve written about this as like what does the maximalist version of a world model look like is you’re incorporating more and more modalities from the universe And you’re training a model on different scales of observations as well.

Swyx [01:30:28]: X-rays.

Anastasis [01:30:29]: And so, yeah,

Vibhu [01:30:30]: You got a good essay that people should read on

Anastasis [01:30:33]: Yeah. Yeah

Vibhu [01:30:33]: Real-world.

Swyx [01:30:33]: No, Meta released a model that was, like, six modalities in one, right?

Vibhu [01:30:37]: Yeah.

Swyx [01:30:37]: I forget what the name of the thing was, but it was like, yeah, okay, depth is one of them, but depth is like a transformation of RGB in some sense.

Vibhu [01:30:45]: ImageBind.

Swyx [01:30:45]: ImageBind, yeah.

Vibhu [01:30:45]: Yeah.

Swyx [01:30:46]: What other modalities? They had heat?

Vibhu [01:30:47]: Audio, depth, heat, text,

Swyx [01:30:51]: Whatever IMU is.

Swyx [01:30:52]: I do think, like, you might as well do ultraviolet. You might as well do, like, just whatever other modality you feel like, ‘cause it’s all data to the model.

Anastasis [01:31:01]: Yeah. And, a big bet is also that there is transfer between all those modalities.

Swyx [01:31:05]: Yeah. Yeah.

Anastasis [01:31:05]: So one of my favorite, examples, which is quite old at this point, is there was this fine-tune of, Stable Diffusion that was called Riffusion Which was

Swyx [01:31:14]: The music one. Yeah.

Anastasis [01:31:15]: Yeah, just fine-tuning, Stable Diffusion on spectrograms.

Swyx [01:31:18]: Spectrograms.

Anastasis [01:31:18]: And it became a quite capable music generator. Right? So there is probably Spatial patterns, so like spatial-temporal patterns if we’re talking about video that emerge at different scales and different modalities. And so there is some degree of, meta-learning that the model has done that allows it to learn faster if you start from a just a model trained on images and train it to predict audio than if you train from scratch on just audio. and there is some other interesting examples. So there is this project called The Well. It’s, it’s a dataset of physics and numerical simulations in physics and biology and a bunch of other domains. So it’s, so it’s essentially different physical systems across very different scales of space and time, from like astrophysics to low-level like atomistic interactions. And we’ve seen. we’ve done some work on this, and we’ve seen that we can take our video model where, real-world video looks nothing like this, and you can fine-tune it on those numerical simulations and just treat them as RGB frames. And you get reasonable performance much quicker than if you just train from scratch.

Anastasis [01:32:36]: Yeah.

Vibhu [01:32:36]: I think we’ve seen this across languages where

Swyx [01:32:38]: Yeah, DeepSeek-OCR as well.

Vibhu [01:32:40]: Yeah, DeepSeek-OCR.

Swyx [01:32:41]: Like, you don’t have to tokenize text. Like, you can just throw them in as images.

Vibhu [01:32:44]: There’s a lot that happens in that base pre-training. Like, there was an argument a long time ago of people saying, “Oh, humans have so many, sensory representations, right? Smell, touch.” Models have a whole two more modalities that we’ll like, that we don’t even have data for. And it’s like, okay, you take AQI sensor, like you can try this stuff, but there’s so much happening in just the base trainer on that you don’t get as much from these little things.

Scientific Data, Cross-Modal Transfer, and Omni Models

Anastasis [01:33:10]: Yeah, exactly. And I think that’s what it solves is data scarcity.

Vibhu [01:33:13]: Yeah.

Anastasis [01:33:13]: So you don’t have as much. You have so much video data available, but you don’t have, like olfactory data that

Vibhu [01:33:21]: The cool thing is it goes the other way too, right? So if you wanna do physics, like if you wanna measure this or you wanna have a diffusion model do audio, it transfers really well. So like in your case, the little bit of post-training for robotics gets a video model to use its fundamentals in another domain. So we can apply that to other stuff too.

Anastasis [01:33:40]: Yeah. And if we look at, like how do you make those models more useful for in scientific domains, and if you look at AlphaFold, they’ve had all these very. Because of the data, the limited amount of data that it needs to be trained on, it’s it’s very fine-tuned architecture just to solve, protein structure prediction. But if you take all those disparate sources of scientific data and you bring them together under a single model, like I think that’s an approach that can help us solve new kinds of problems across science by leveraging all the learnings from one modality or one set of, data to another. So very early days for that direction, but I do think that’s where ultimately what the end game of simulating the world is. You’re not just using RGB. You’re using RGB as a starting point, but you can incorporate more and more modalities of the universe and leverage the transfer that happens from learning from one to the other.

Vibhu [01:34:42]: I guess the follow-up there is what’s the drawback of omni? Like, why is everything not an omni model? Also, why not now, and why. Would you start from language backbone or image video backbone and then go omni from there? Does it matter?

Anastasis [01:34:57]: Yeah. We need to take it one step. We need to solve robotics first, and then we can go into

Swyx [01:35:02]: Solve everything now.

Anastasis [01:35:05]: Yeah. I do think there is a lot of open-ended research that needs to happen for, those omni models. There is a lot of things that require careful consideration when you’re bringing multiple modalities into a single model to predict. But I think, I expect those to be solvable.

Closing: Film Festivals and the Future of AI Video

Swyx [01:35:23]: Wonderful. you’ve been very generous with your time. Congrats on all your success, and, yeah, I’m excited for the, AI Summit, or physical AI Summit.

Anastasis [01:35:32]: Yeah, thanks for having me.

Swyx [01:35:33]: And yeah, and people should check out the film festival if it’s in town, right?

Anastasis [01:35:37]: Yeah.

Swyx [01:35:37]: Yeah. You’ll be gonna be touring all over the place.

Anastasis [01:35:39]: Yeah. Next year we’re probably gonna do that. So we do film festivals every May or June of

Swyx [01:35:45]: Yeah.

Anastasis [01:35:45]: And we did the last one in New York, LA, Tokyo, and at the AI Engineer,

Swyx [01:35:52]: Yeah

Anastasis [01:35:53]: Fair.

Swyx [01:35:53]: Yeah. Yeah.

Anastasis [01:35:54]: So yeah, hopefully even more places next year.

Swyx [01:35:57]: No, I think like someday, you will be hosting the Oscars of AI video, and, I think people should like take this very seriously as like a potential career they can have.

Anastasis [01:36:07]: The Oscars of AI video will be called the Oscars.

Swyx [01:36:10]: All right. All right. Thank you.

Anastasis [01:36:14]: Thank you.

💾

  •  

Can Skills Learned in Games Transfer to Real-World Work?

“Games have always been these underrated educational tools. They’re super approachable. They’re very human.”

Those are the words of Alex Duffy, co-founder and CEO of Good Start Labs, who spoke to Latent Space about why his company is turning games into training material for AI models. The company was spun out of AI media and tools company Every last October, with $3.6 million in funding from General Catalyst, Inovia, Every, and angel investors.

The idea came from a 2025 Twitch stream of frontier models playing the game Diplomacy, which Duffy said normally takes “days or weeks to play.” This was when he worked at Every as its head of AI training.

2025 Twitch stream showing a Diplomacy betrayal by OpenAI’s o3 model.

(For more on Diplomacy and LLMs, see our interview last year with Noam Brown, soon after he won the 2025 World Diplomacy Championship!)

Watching the AI agents battle it out in Diplomacy showed Alex Duffy how each frontier model acts differently when faced with gaming scenarios. In particular, he noticed the OpenAI model (o3) winning all the games by planning a future betrayal, whereas the Claude model (Opus 4) refused to lie and thus “got destroyed.”

From this, Duffy concluded that training AI models on games like Diplomacy could teach them skills like strategic thinking. Especially because those kinds of games have outcomes that can be verified. In a later article published on Every, Duffy wrote that “fine-tuning a model on the strategy game Diplomacy improved its performance on customer support and industrial operations benchmarks.”

Grok 4 Fast is the least likely to betray you in Diplomacy, according to these September 2026 rankings by Good Start Labs.
On the other hand, don’t trust Gemini 2.5 Pro in Diplomacy!

Duffy and his co-founder Tyler Marques launched Good Start Labs with the intention of exploring other games that could teach useful skills to AI models.

“It became really clear that reinforcement learning environments were one of the most reliable ways to teach models anything you could verify,” he said.

The bigger idea is that the way a game is presented to an AI can determine which skills it learns, and whether those skills carry into work outside the game. And the best evidence for that so far comes from a nineteenth-century railroad game.

When game training transfers to financial research

Good Start Labs recently trained a 30B model inside the game 1830: The Game of Railroads and Robber Barons, described on Wikipedia as “a strategy game where the only element of luck involved is in determining the initial play order.”

They then tested the same model on financial research tasks. The experiment was designed to test whether habits learned in a game could transfer outside the game.

“That game has a stock market mechanic within it,” Duffy explained. “You’re bidding on stock of these railroad companies to try and create this logistics network. And we’ve set up tasks where models are going through a database to find information about how the game’s been played, putting it into an Excel file, reasoning over it, creating some functions within it, and then calculating its answer in that way. And so it mirrors what you would typically do in a finance workflow, but you’re doing it in this game.”

The published results compare single-turn question answering — where the model is presented with a game state and asked to make the next move — with a “multi-turn terminal agent that uses tools to explore its environment, plan a strategy, and adapt in real time.”

Both training designs improved their respective in-game objectives, but only the terminal-agent design improved performance on the Finance-Agent benchmark.

Designing a learning environment to teach capabilities

The 1830 result showed that the training design is key. But more generally, Duffy said Good Start Labs can also add an expert model that provides denser, stepwise rewards.

The harness also allows an environment to approach the same game in different ways.

“How you design that [the harness] totally changes what the model can learn,” Duffy said. “You can imagine a model that is looking at pictures is going to learn different things than one that’s reading through natural text [or] one that has everything framed as Python.”

Duffy described the overall goal of Good Start Labs as figuring out “how do you design a learning environment to teach specific capabilities?”

That question is explored in COS-PLAY: Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks, a paper co-authored by Duffy and Marques with researchers from several universities.

In the paper, the system gives a decision agent access to what the authors call “a learnable skill bank to guide action taking.” A separate skill-bank agent studies the trajectory and makes changes to the skill bank, which is then looped back for the next run.

What about the latest frontier models?

I asked whether Good Start Labs has compared newer models, such as Claude Fable 5.1 and GPT-6 Astra, in the same game environments? And by extension, do increasingly capable base models make the harness and training environment less important?

“We compare every new model,” Duffy replied, adding that the newer, more capable models tend to be better at the games. However, similar to what the original Twitch streams showed with the 2025 models, the new models “diverge on the personality axes: betrayal, collaboration, theory of mind, etc.”

As for harnesses, he said that “a more capable model needs less handholding to finish the same task, certainly.”

But for what Good Start Labs is doing — treating “the environment as curriculum” — the harness “matters more, not less.”

“GPT-6 Astra reports doing less chain-of-thought and jumps to answers,” Duffy said. “If you want a model to work a certain way while solving a problem, the harness is what forces it. Astra can probably do the math in its head, but you’d rather it use code so you can trust the result.”

What is Good Start Labs selling?

In a recent blog post, the company described its work on “improvement loops,” which include training systems, harnesses, and observability. But how does that translate into products that Good Start Labs offers other companies?

“The main thing that we sell in terms of AI improvement is data and learning environments,” Duffy replied. Its main customers are frontier labs — for which they provide reinforcement learning data to help further train their models.

He describes the data part of its offering as one of two things. The first is “trajectories of agents playing games” — what an agent observed, what it decided, which actions it took and what happened afterward.

The second is custom data for specific game publishers, where “agents are live in their games.” The agents can play inside those games, generating interactions that may be useful for training and evaluation. Duffy said any data sold to model developers is anonymized and stripped of personally identifiable information.

The learning environments that Good Start Labs sells are “full games where models can play end to end,” he noted. It isn’t about winning the games, though. It’s more about teaching AI models to solve problems.

“We’ll also make a lot of tasks where the models are using the game engine as the verifiable source of rewards, but are solving problems in a way that you might not expect.”

So can game skills be transferred to real-world work?

Alongside its custom work for clients, Good Start Labs is also training a general model from the expert models it has built for specific games. The idea, said Duffy, is to unify those expert models “into this general game intelligence that could be applicable everywhere.”

But while the 1830 experiment suggests that agentic game training can transfer to a structurally similar financial-research task, it’s unclear if there will be broader real-world transfer. I asked Duffy what the evidence actually states today?

“Today’s evidence supports pretty clearly that goal-directed execution matters, and reasoning transfers,” he replied. He pointed to a recent article by Surge AI showing that office work post-training improved coding, adding that “DeepSeek R1 showed it more broadly.”

“We’ve seen it twice ourselves: the 1830 finance task, and Diplomacy training that produced a better customer support agent. Every environment we’ve built also improves tool use downstream.”

So that makes the answer to our big question a qualified yes: some game skills can transfer to real-world work. But Duffy says how broadly and reliably they transfer remains an open question.

  •  

PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors

GitHub invented pull requests, and for 18 years they have been open by default. But now some of the top AI-native open source projects are shutting PRs off, because they’ve found a better way.

These projects, which include Flue and tldraw, refuse to accept PRs from external contributors — in part because they’re usually AI-generated. Instead, the maintainers prefer to use their own agents to create and manage PRs.

Also, many projects have begun using a “software factory” to manage community contributions. Typically this involves a ‘team’ of agents triaging a PR, reproducing the issue (if it’s a bug), implementing a fix or a new feature, reviewing it, and then handing it back to a human to merge it.

Vercel’s software factory for AI SDK

Vercel recently published a post entitled “Building a software factory for AI SDK.” It describes how the open source AI SDK project, which gets over 20 million npm downloads per week, deployed agents to get control over its PR and issue backlog — which had reached “over 1,000 open issues and almost 800 pull requests” by late June.

There are several types of agents in Vercel’s system, each of which focuses on a different task. For example, there’s an agent that reproduces a bug, another that applies a fix, and yet another that reviews the fix.

Diagram from Vercel; comments by Latent Space

One of the key reasons why Vercel set up this software factory is because it trusts its own agents to do the work, more so than agents run by community members.

“If we have a very specific agent with a very specific prompt that we optimized — and we know that, over history, it was very successful in fixing a certain category of bugs — then we develop trust in that particular agent configuration,” Vercel engineer Lars Grammel explained in a YouTube video.

“For open-source projects, it’s worth considering having your own agents and your own setup, and not necessarily trusting the community, because it can actually cut down your time to review,” he added.

Example of software factory workflow in AI SDK project.

Grammel also showed the deployment architecture for its system, noting that “there is a UI, there’s a web app, there’s an underlying API, there’s an execution space, and there are sandboxes.” It’s then synchronized with GitHub, which automatically triggers other actions. The UI Grammel mentioned was custom-made.

Vercel’s software factory deployment architecture; diagram by Lars Grammel.

Just four weeks after this software factory was implemented, Vercel claims the factory now “authors between 25 and 35% of PRs we merge and closes 70-80% of issues.”

Astro’s auto-triage system

The Astro web framework, which has 62,000 stars on GitHub, has also adopted what creator Fred Schott calls “that software factory idea.”

“For five years, we were in this place where issues came in faster than we could handle them,” Schott told Latent Space.

But now, with agents handling the triage work, they’ve reestablished control.

“It’s totally shifted in the last six months,” he said. “We can now solve these issues with these automations — handling triage, reproduction, getting the user to actually verify the fix that the bot is suggesting before we even look at it.”

Example of an Astro factory bot in action

The result was not just a large decrease in open issues, but a complete change in how the Astro team deals with incoming community requests.

“I’ve never seen that in my entire decade-plus experience with open source,” Schott said. “Being able to essentially treat issues as a thing that every week, you prioritize — no matter what — versus a backlog that you’re constantly trimming.”

Furthermore, the Astro “auto-triage” system directly led to Schott creating a brand new agent framework, called Flue.

Flue doesn’t accept your PRs, but is open for discussion

With Flue, Schott is trying an even more radical approach to PRs. Flue’s contributor guide states that “we’re going to try to reimagine things” — partly to prevent what it calls “Drive-by AI slop PRs.”

Basically, Schott explained, every external pull request in the Flue project is automatically closed and converted into an issue or discussion. Bug reports and fix proposals get turned into issues, feature requests become discussions.

Agents can do most PR tasks now, according to Flue’s contributor guide.

“If you submit a PR, no hard feelings, we’re just going to go and represent it for you as issues and discussions. And from there, trying to figure out the right way to bring people on.”

It’s kind of like treating incoming requests as leads, rather than as a piece of work a maintainer feels obliged to review. The contributor guide explains that it uses the team’s own expertise combined with “the best available SOTA [State-of-the-Art] LLMs that we have access to” in order to help them decide what to work on next.

Once a decision is made in the issue or discussion, agents are then deployed for “research, design, implementation, and initial review.”

If our agents write the code, your external PRs are worthless

Like Flue, the “source available” React drawing tool tldraw (50,000 stars) automatically closes external PRs.

Project creator Steve Ruiz announced this policy in January and five months later reiterated it, noting that it was “an opinionated decision made in response to changes in how we’re coding (more discussion, more agents), the social practices around public contribution, and the changing landscape around code security.”

HashiCorp co-founder and Ghostty creator Mitchell Hashimoto, now a co-founder of Superlogical, takes it even further. He thinks “the future is that large open source projects will close contributions completely.”

Ruiz responded, “It just makes less sense to have people contributing code if the issue is decently well-specified and the code can be written by agents.”

But…what happens to the community?

Traditionally in open source, pull requests have been reviewed by maintainers not only for the code, but to teach contributors and assess them as future maintainers. If projects like AI SDK and Astro are using their agents to do much of the code review and implementation, where does that leave community members who want to be more actively involved?

Schott recognizes this as a risk.

“It still leaves this open hole of, well, if you just keep narrowing the project, at a certain point, you and I go on vacation — what happens? It doesn’t really solve every problem.”

However, the fact that both Flue and tldraw don’t accept PRs but do accept new issues and discussions perhaps points to a solution. Which is that by talking to each other more, community members better get to know — and trust — one another, which is both a way to learn from peers and potentially prove yourself worthy of being a maintainer.

Example of a tldraw issue (above) being turned into a PR (below)

As for the code, if it’s easier for maintainers to use AI themselves than to accept external code contributions, then as tldraw founder Steve Ruiz put it, “it’s better to limit community contribution to the places it still matters: reporting, discussion, perspective, and care.”

  •  

Lovable CTO: The Future of SaaS Is Apps That Agents Can Use

Lovable is well known as an AI-powered platform to build applications. But ironically, it is now moving towards a future where fewer and fewer people will be using conventional apps. That is, of course, because of the growing impact of agents.

In a recent blog post, Lovable outlined a vision for “a digital brain for your team connecting your daily tools.”

Or as Lovable CTO Fabian Hedin put it in an interview with Latent Space, “you can get to a place where you’re using one entry point to all the work that you’re doing.”

Diagram by Latent Space based on an internal diagram shown to us by Lovable.

To be clear, Lovable still wants to be the tool you use to build apps — but increasingly, it will also enable you to build what Hedin calls “capabilities.” Lovable defines a capability as a useful part of an application that an agent can call directly; bypassing the need for a human user to open the app.

Lovable can turn a published application into agent-accessible capabilities by exposing selected functions from the app as tools through a hosted MCP server. The result is essentially one application with two interfaces: a traditional human UI and a new agent interface that can be used from ChatGPT, Claude and other MCP-compatible AI clients.

Diagram supplied by Lovable

This is how fast an AI business evolves

This shift towards capabilities is the latest evolution from Lovable, in an already fast-moving 3 years in business.

Lovable emerged from GPT Engineer, an open source coding tool that launched in 2023, initially focusing on prototyping. In November 2024, it became a commercial product and the following month, it was rebranded as Lovable.

By that point, they’d begun to notice some of its users building production apps on Lovable — including products that had become real businesses.

“We started seeing people on the platform building not only a prototype, and not only an MVP [Minimum Viable Product], but the actual thing — an actual product that serves real customers,” Hedin said.

Next, Lovable noticed its users creating internal software, in some cases to support a public-facing app and in other cases as an internal app built for an enterprise company.

“People started creating not only software to enable a business, in terms of a customer-facing product, but also the operations behind the company,” Hedin said.

He means tools like a CRM, an admin panel, or a customer-support console.

From app builder to agent platform

So in less than three years, Lovable has become an all-round software creation and hosting company, which means it’s swimming in the same waters as the likes of Vercel and Cloudflare. That said, Lovable is more focused on AI-generated software than infrastructure. But we are seeing crossover in these markets — for example, Vercel’s v0 allows you to generate an app from natural language, just like Lovable.

Also just like the black triangle and orange cloud companies, Lovable has expanded into agentic workflows.

Lovable connectors, which let you use external tools.

This rapid product evolution has been accompanied by strong user and revenue growth. According to a tweet from Deedy Das, a partner at lead investor Menlo Ventures, the company has surpassed a $500 million annualized revenue run rate, with more than 60 million projects created and over 900 million monthly visits to Lovable-built apps. Lovable also says employees at nearly two-thirds of the Fortune 500 have used the platform.

Unsurprisingly, Menlo Ventures is doubling down on its investment. It led Lovable’s $400 million Series C this month, alongside the Scaleup Europe Fund managed by EQT, valuing the company at $13.3 billion.

Hedin attributes the pace of change to a combination of Lovable’s innovation and the rapidly improving state of LLMs.

“Every few months, we introduce new capabilities at the application layer, while the large language models also continue improving. Those two things compound.”

Lovable’s model of a company brain

The concept of a digital brain for an organization, for Lovable, essentially means a single interface where you can access many different tools and workflows.

“It should have as much context as possible about you, your company and the world around you,” said Hedin. “Then it needs the capabilities to perform both general tasks and actions that are specific to your organization.”

Ultimately, he added, the goal is that “everything that you’re building can be reused in an agentic way.”

Diagram supplied by Lovable

In a sense, then, applications are becoming a collection of capabilities that users will increasingly access through an organizational agent — instead of, or in addition to, the actual application.

“Our job as a platform is to ensure that all these separate capabilities are connected through one agent — not that you have to build a different agent for every task,” said Hedin.

As an example, Hedin mentioned an internal application they use at Lovable.

“We built this internal tool to help us grant credits to users [via] our support team, and help manage our platform in different ways. Those capabilities are now available [internally] through the Lovable agent.”

Lovable also wants this company brain to work asynchronously. Its agent can schedule itself to resume a task later — for example to check a deployment or to monitor a recurring process — then return the result to the same conversation.

The competition

Lovable isn’t the only company pursuing a “company brain” vision. Vercel CEO Guillermo Rauch recently introduced its internal agent, called @𝚟. “Every day-to-day job at Vercel now involves @𝚟,” Rauch tweeted. “It’s growing exponentially both in daily interactions and token use.”

Hedin acknowledged that Vercel and other AI companies are building towards a similar vision, but he thinks Lovable’s “wedge” is “being the best place to build the capabilities that agents need.” In other words, Lovable’s focus is on helping their users build the capabilities that a company brain will need.

“Orchestrating these capabilities is the easy part,” Hedin said. “Making sure they are well connected, built correctly and reliable is the hard part.”

He also hinted at why they’re using the word ‘brain’ to describe this shift, rather than just ‘agent’.

“I’m careful about using the word ‘agent.’ It suggests something like an employee performing a task, which is an easy way to think about it. But underneath, it is really about connecting the right context and capabilities.”

Security and connecting to external capabilities

Perhaps the biggest challenge with the agents and capabilities paradigm is security. For instance, if one of your employees creates an app with Lovable that connects to the company Slack, you want to ensure that user doesn’t inadvertently expose their personal messages, or any other confidential information, to the company brain.

Connectors are Lovable’s method of connecting to external tools and services. Hedin said the platform must account for a “kind of permissioning graph” to maintain security and privacy.

As described in a technical article on Lovable’s blog, one connector type, which Lovable calls an “app user connector,” preserves each user’s identity and source-system permissions. Credentials are stored server-side in encrypted form and handled by Lovable’s connector gateway, rather than being exposed to the generated application; the app instead presents a short-lived key bound to the relevant user.

Diagram supplied by Lovable

“We separate the connection to external systems from the application code being written,” is how Hedin put it. “The application interfaces with the Lovable platform, but the application itself never gets access to those credentials.”

The future of SaaS

So Lovable is moving to a future where a company brain uses capabilities derived from the apps its users build. That begs the question: what will happen to SaaS apps?

Hedin reiterated that people will increasingly interact with software through an AI layer — the company brain concept.

“People are not going to have as many tabs open in different tools as they have historically. That experience is going to consolidate, but the vertical capabilities those tools provide will remain valuable.”

He recognizes that some traditional SaaS products may “fight” this trend, by sticking with their traditional apps and not adapting, but he says Lovable wants to become a platform for building capabilities.

“We want to build this open platform that anyone can connect to, anyone can use,” he said.

Hedin ended with some advice for SaaS companies, whether existing ones or apps that might emerge on the Lovable platform.

“I think SaaS businesses are going to have to focus more on providing the shovel for AI to use their capabilities.”

  •  

The /wayfinder Skill: Navigating the “Fog of War” of Planning

We’re currently developing a new series about skills, with the aim of giving you a regular supply of new skills to use in your projects. We’re kicking things off with an interview — and a super-useful skill — featuring Matt Pocock, whose “AI Skills for Real Engineers” project has over 220,000 stars on GitHub. He also talks about these skills to 347,000 subscribers on his YouTube channel.

Pocock recently released a new skill called /wayfinder. Its purpose is to help you and your agent figure out a project where the end state isn’t entirely clear. Or as Pocock put it in our interview, /wayfinder helps you navigate “the fog of war,” where you have a project but “you can’t quite decide everything right at the start.”

The following interview has been slightly condensed for readability, so you can read it, absorb Matt’s insights, and then test out /wayfinder for yourself!

Latent Space: What were the goals of wayfinder?

Pocock: What I noticed is I was doing a lot of work with AFK agents [Away From Keyboard] and trying to schedule in a ton of work so that my agents could run virtually overnight. I would just plan a bunch of stuff, and then I would create a spec and then turn that spec into tickets. And I had a really well-developed set of skills for how to turn work into scheduled stuff that agents could just crack on.

But [...] I was finding the planning stage really onerous, because I would have to be constantly thinking about my session management. Like, how many tokens am I into my context window? How deep am I going here?

Matt Pocock’s wayfinder skill, as documented in GitHub

I didn’t want to feel constrained in the planning stage anymore. I wanted an orchestrator layer that would basically say, okay, whatever you want to plan, I’m going to handle the planning sessions for you. I’m going to split this out into multiple different threads, do prototyping, do research and pull it all back together, so that you don’t feel constrained in the planning anymore.

And then your specs can be even more detailed, and you can just whack off an AFK agent to go and do tons more work.

Latent Space: What was the design process of coming up with this skill?

Pocock: I had this kernel of an idea of, what if I didn’t have to manage the handoffs? What would that look like? And then, what would it look like to have some kind of centralized document to have all of those pieces together?

Whenever you’re thinking about context management — because that’s really what a skill is, you’re managing the context of the agent you’re working in — you need to think about the information flow. So what I wanted to think about is, what if a grilling session could manage other grilling sessions? What would that look like?

Well, the first step to that is, what does the grilling session that’s being managed need? What does the child need in that situation? So the child probably needs to understand a vague overview of what else is happening, and they need their specific task.

So there, you’ve got two documents. You’ve got a map — which is all of the rest of the stuff, all the decisions that have already been made. And then you’ve got the specific ticket that goes into the actual session. And what you notice there is that those words are very precise.

You’ve got the map, and you’ve got the ticket, and you’ve got the session. And once you’ve got the kernel of an idea, you then need to come up with the words for that idea. Because once you’ve figured out the words, then those entities can be really clearly mapped out by the agent.

Because if you just call everything a ticket, or if you just refer to it in different ways in different places, then it’s going to be really confused and you’re going to get strange behavior. Whereas if you use these very specific, what I call leading words, to lead the agent to understand exactly what each part is, and you’ve understood what the information flow is, then you’ve got your skill.

Latent Space: What kind of use cases do you think wayfinder would be useful for?

Pocock: Well, I’ve been using it for all sorts of stuff. I’ve been using it to actually plan courses as well. In wayfinder, there are different types of tickets.

So you’ve got grilling tickets, which are just a grilling session. Then you’ve got prototype tickets for creating prototypes, research tickets for creating [and doing] research, and then task tickets — which are really broad…basically, just anything the human needs to do that the agent can’t do. And so once you think about that, you realize, OK, I can apply that to anything.

One really key idea in wayfinder is the ‘fog of war’. So this is the concept of, you can’t quite decide everything right at the start.

I decided to test /wayfinder on a project to rearchitect my personal website. Here’s the initial project set-up, in this case using Claude Code.

You can make certain decisions, and those certain decisions sort of lead you there and push further out into the fog of war — kind of like Warcraft III style, exploring the map. And once I had the idea of ‘fog of war’ and ‘map’, I realized those two terms actually work really nicely together, and it really leads the agent into the right idea. So I’ve been using it for engineering, for non-engineering stuff, for course planning, all sorts.

Latent Space: This concept of the fog of war — it’s weird to consider what you don’t know that you don’t know. Maybe LLMs are good at capturing that.

Pocock: I feel like with the grilling stuff that I’m still working on, that captures an idea that you don’t know stuff, but maybe the agent can contribute something and illuminate a part of the room that you don’t quite understand yet.

And wayfinder is just sort of an extra layer on top of that.

Working through my website rearchitecture project using /wayfinder. There’s 20+ years of content to re-organize!

Latent Space: Yeah, and there’s all these artifacts. How much time do you spend teaching the model all this terminology?

Pocock: For the last few months, I’ve been pretty obsessed with terminology — and finding the right terms for certain things. I’ve put together, I haven’t actually put it out yet, but it’s an AI coding dictionary — of basically all the terms in AI coding. It’s in this beautiful graph that you can explore and understand exactly what an agent is, exactly what a harness is, exactly what a model is, blah blah blah.

I’ve redone all my courses to use that dictionary and make it really solid. And then all of my skills use a consistent dictionary as well. So they’re all working off the [same] assumptions, the same leading words.

I realized that I needed a ubiquitous language between me and the agent. Between me and the agent, there is a communication barrier. And that’s what I’m trying to do with my skills all the time, is try to find the right words.

And agents are really good at showing you the opportunities for different wording — really good at domain modeling, actually.

Latent Space: When do we directly use the grill-me skill, versus wayfinder?

Pocock: Use ‘grill me’ in cases where you feel like you can plan the whole thing in a single session, and you need to align before you go. So most small features will fit into this. Most stuff where you can see the path ahead of you, but you just want to make sure the agent is on board, ‘grill me’ will work with that.

For stuff where you don’t know the path ahead, for stuff where you can feel the fog of war in front of you, use wayfinder. You’re gonna find your way with wayfinder. So that’s how it works.

  •  

Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing

With the intense competition among frontier model companies, together with ever-increasing power of open-weight models like Kimi K3 and Qwen3.8-Max, model routing has become a key part of AI deployment. We’ve just seen Stripe buy OpenRouter for over $7B, but the trend is equally hot in enterprises.

Glean, co-founded and led by ex-Google Distinguished Engineer Arvind Jain, specializes in bringing AI to large organizations. It was last valued at $7.2B after a $150M Series F fund raise last June. This year, it reached $300 million in annual recurring revenue (ARR) — a three-fold increase over 15 months.

Part of Glean’s mission is to select which model to use for each task — or indeed if an LLM is even required.

“A big goal of Glean is to avoid using LLMs for tasks where we don’t need them,” Jain told Latent Space. “Sometimes you’ll see queries in Glean where people are adding two numbers or multiplying two numbers. They could have used a calculator to do that.”

But what Glean is mostly trying to do is bring what Jain calls “one really powerful personal co-worker” to enterprise employees. And that means being a kind of meta-harness for leading LLMs.

Glean announced its third-generation Glean Assistant last September; these days, agents are a big part of Glean’s system.

“You can think of Glean today as a superset of ChatGPT, Claude, Gemini, Grok,” Jain said. “All these different AI products that we’ve been using day to day, Glean combines the power of all of them into one experience.”

With enterprises, bringing AI technology into an organization is just half the challenge. The other half is bringing organizational knowledge into the AI systems.

“Ultimately our business is to deeply understand your data, knowledge, and information, but also how work happens inside your company,” Jain said.

How model routing is done in Glean

So what does model routing mean in practice? Basically, Glean offers three levels of model selection:

  1. Employees can explicitly choose a model.

  2. Administrators can restrict models or impose usage limits.

  3. Glean’s automatic mode selects a model dynamically for each task.

Configuring models for certain tasks.

It turns out automatic mode is mostly chosen by Glean’s customers for economic reasons.

“Why are people talking about model routing? Why are they excited about it? It’s mostly because of cost,” Jain told us.

Another co-founder of Glean, engineering lead Tony Gentilcore, recently claimed that Glean “is 4x more cost-effective” than Claude Code, “averaging $0.45 per task versus $1.84 for Claude Cowork.” He put that down to Glean’s “harness and routing capabilities.”

Individually, many of us are getting great value out of our $20, $100 or $200 monthly subscription to an LLM provider. But for an enterprise, the per-user costs can easily spiral out of control.

“AI models have been getting expensive,” Jain said. “Like, if you look at Opus or the latest models of GPT, the most advanced models. Not only are they very powerful, they can run much more complex tasks than the previous models. But on a per token basis, they’re more expensive — sometimes double or quadruple the rates of the previous models. And then users actually use them to run much longer tasks. So you’re spending, like, 10 times, 20 times, more, on a per user basis, than what you were doing last year. So the costs have gone up a lot.”

The human feedback loop

Another key factor in Glean’s rise is that it gets to see how ordinary business users are using AI. The product is potentially deployed to every employee as a “coworker,” and it’s also used to build and deploy agents across all departments and functions.

Among its customers, Zillow reports 80% adoption across 7,000 employees, while at Booking.com, “Glean became the first AI platform adopted company-wide.” That kind of penetration gives Glean an enviable view into how AI is being used in enterprises.

“So we are getting to observe what people are actually doing with AI on a very broad basis,” said Jain. “We are getting to see when they’re on different types of tasks with AI, what models do they select first, and when they are not satisfied, when they actually upgrade to some other model [that] actually gives them the right results.”

This human feedback loop, at scale, helps improve the model routing system.

Here’s Waldo, gathering raw materials

Another part of Glean’s architecture is a model called Waldo, which Jain described as sitting on top of the large language models. Waldo was introduced in April as “Glean’s first agentic search model.”

Glean claims that Waldo, its agentic search model, “reduces latency by 50% and tokens by 25%, reserving advanced models for work that needs them.”

In a technical blog post, Waldo was portrayed as a kind of filtering process for user queries: it “decides how to break down the question, which tools to use, what to read next, and when it has enough evidence to hand off to a frontier model for a high-quality answer.”

This means the model routing is happening after Glean has determined what Jain calls the “raw materials” that are needed for the task.

“We’re able to assemble the raw materials needed to do the work without burning LLM tokens,” he added.

A corollary of this is that a cheaper model with better context may outperform a frontier model loaded with irrelevant data.

The rapid rise of open-weight models

Jain confirmed there is now significant interest from enterprises in open-weight models, primarily due to cost concerns. But this has only happened over the past few months.

“Last year, the usage [of open source LLMs] was minuscule and nobody was really seriously considering open source,” he said. Partly that was because of the “stigma” of many of these open source models being developed outside the US.

But suddenly, interest among enterprise customers has risen.

Jain’s tweet on July 27, 2026, in support of open-weight models.

“So in the last three months, because AI got so expensive, businesses have started to find it untenable to maintain these AI investments,” Jain said. “Given that open source is an order of magnitude cheaper to do tasks, it has created a lot of interest. Today, I can say that in most enterprises, they are considering open source models to be a key part of their AI strategy.”

More than that, organizations tend not to rely on just one or two providers anymore — and the rise of open-weight models is driving this trend.

“Nobody is willing anymore to rely on only one model provider, or two, and nobody thinks that they can survive without open source,” Jain said.

Evals

You can’t have a serious conversation about AI in 2026 without discussing evals — assessing the quality of results from LLMs. I asked how Glean goes about doing evals and how that is fed back into the model routing system.

Jain said they have “internal testing systems” where they compare real-world workloads, across different query classes, with alternative options. So they let the model choose a route and in parallel they try to complete the same task with “some other models which are maybe a little bit less expensive and a little bit more expensive.”

How Glean monitors quality.

Glean then uses “AI-based judges” to determine “how spot-on the model router was.”

“So there’s this continuous learning that gets updated with new real-world traffic, where basically what is happening is that you let the model router do the work for the user, but behind the scenes you run the same task,” Jain explained.

He added that this is done for only “a small fraction” of the real-world usage, but at Glean’s scale that’s more than enough to help train and improve the model router.

From enterprise search to end-to-end AI platform

One of the trends we’ll be monitoring going forward on Latent Space is how AI systems are being implemented within enterprises — and how some of these organizations are going full-on AI-native.

Glean is an especially interesting company to monitor for these trends, since it was one of the very first enterprise-facing AI companies. It was founded in early 2019, initially to tackle enterprise search. As Jain put it, Glean was “the first player to work with transformers and language models for businesses.”

Glean’s AI Answers draws “directly from your organization’s documentation.”

In April 2023, swyx interviewed Deedy Das of Glean. Das, who is now a partner at venture firm Menlo Ventures, was a founding engineer at Glean. But even at that point, in 2023 — about four years into Glean — the focus was still mostly on enterprise search.

Now, in 2026, enterprises aren’t just using AI for search. AI is becoming an integral part of every employee’s workflow.

That makes Glean a much ‘sexier’ AI company, as Das himself said on his return to the Latent Space podcast last November. “Broadly, one of the things that I love about Glean is it’s such a boring unsexy company that became sexy later,” he said.

This brings us full circle back to model routing. Arvind Jain ended our discussion by calling Glean an “end-to-end AI platform” that gets “used very heavily” by its enterprise customers. This, he added, allows Glean to “have that data that is required to do effective model routing.”

  •  

React for Agents: Astro Creator Brings Hooks to his Meta-Harness, Flue

Agent frameworks for developers are still at an early stage, with the likes of Vercel’s eve and Fred Schott’s Flue — both launched this year — setting the early template.

Schott is the creator of the web framework Astro, which led to his company being acquired by Cloudflare in January. He’s just released version 2 of Flue, its first stable release, which has as its foundation React-style “Agent Hooks.”

In Flue, an agent is represented by a JavaScript function. This function “re-renders on every turn,” meaning before every model call.

The addition of hooks came after Schott realized that React’s composability would be a great fit for agent development.

“I originally tweeted that we were building the Astro for agents or the Next.js for agents,” he told us. “But then I realized: maybe no one has even built the React for agents.”

Editor’s Note: we last talked about the React for Agents with Bret Taylor, CEO of Sierra and Chairman of OpenAI:

“We’re still trying to figure out who the reactive agents are and the jury is still out… We’re sort of in the jQuery era of agents, not the react era.”

Hooks are authored in TypeScript. According to the Flue 2 launch post, they “let you build dynamic agents that can manage their own state, listen to agent lifecycle events, and even attach different resources and capabilities dynamically to enhance themselves at runtime.”

There are 16 built-in hooks in Flue 2, including useSkill(), useTool(), useSubagent(). You can also add custom hooks.

How Flue evolved via React-style hooks; diagram by Richard MacManus

What hooks open up for developers is that they make an agent much more dynamic, by allowing its configuration to change as a conversation or workflow progresses. Schott said this is needed to build “real support bots, real triage bots,” because they can’t be fully configured in advance. The agent can’t just be static — it has to adapt in real-time to what the user wants or the situation demands.

Agent hooks bring those capabilities to Flue. For example, a support agent might bring in an account management tool after first verifying a user.

File based magic is an antipattern

Schott’s thinking about how to build an agent framework has evolved rapidly since he publicly launched Flue 1 in early May. Initially, he wanted to take existing web framework concepts and apply them to his new agent framework. He uses file-based routing as an example.

“So we kind of naively ported that over to Flue, thinking — great, well, I’ll put your five agents in these five files, and that’ll be the five routes that they expose. But for a lot of people building with Flue, especially the bigger customers, their whole company is one agent. They don’t care about routing. There’s one agent.”

So after the first Flue users showed these early patterns, composability became front of mind for Schott. That led him back to React.

“As you can see from the Flue 2 API, we’re taking it more from React [...] than we are from Astro or Next.js — where it’s less about routing and these website concepts and more about, at its base level, how do you compose an agent on many different things?”

Flue’s central proposition: agents need a harness

A key concept in Flue is that an agent must have a harness — meaning that it’s in an environment where it has access to the context and capabilities needed to accomplish various tasks.

“Instead of you and your code driving the LLM and telling it what to do with scripts, you’re putting the agent into this harness, and it is able to drive itself and work through problems,” explained Schott.

Flue is built on top of Pi, an open source minimal harness. Essentially, Flue is an opinionated take on Pi — adding features that Schott thinks are helpful to developers building agents. For example: hosted agents in Flue 2 are now built with Vite, an open source build tool.

Indeed, Schott likens Pi’s role to the foundational role that Vite now plays beneath Astro.

“I think Pi can serve that role, where it’s the right abstraction — it doesn’t do too much, but it gives the right APIs that then we can go and say, well, let’s have an opinionated take on this that does more.”

Building on Pi meant committing to having a built-in agent harness.

“Our early bet was that the harness is actually not a feature, but it’s fundamental to what you think an agent is,” Schott said. “There is no agent without a harness.”

Building Flue agents with coding agents

The Flue project began earlier this year within the Astro repository, as an issue-triage system. At first, it was an LLM-driven script or workflow reviewing issues. But then, explained Schott, it gained the ability to take actions in the repo.

“It started to transition from just automation in a repo to wanting to take the Claude Code experience, make it headless, make it hostable and run it in the cloud.”

So that’s when the idea of a harness as anchor emerged. Indeed, in his v1 launch post in early May, Schott described Flue as “like Claude Code, but 100% headless and programmable.”

I myself tested out Flue using Claude Code, which guided me through setting up my first Flue agent. And Schott confirmed this is how many developers use Flue.

“We very much are building for them,” he said, regarding AI coding agents. “Our whole onboarding flow is that, you know, pass this prompt to your agent, it’s gonna guide you through it. All of our docs have markdown support.”

Where Flue fits in the agent development stack

The closest comparison to Flue is Vercel’s eve, which also treats the harness as foundational. Vercel and Cloudflare have been known to beef in public, but Schott is generous in his opinion of eve.

“Eve, I think, is the most directly competitive,” Schott said. “It came around at the same time, so it had that same take that a harness is built-in.”

Schott also referenced what he called the “OG agent frameworks,” which came before Flue and so weren’t created with a harness as the central concept. He listed Vercel’s AI SDK, Cloudflare’s Agents SDK, and Mastra (developed by the same team that built Gatsby, a web framework predating Astro).

While these “OG agent frameworks” are all adding harnesses now, Schott considers that an added feature — whereas Flue and eve both have built-in harnesses.

I asked where Flue sits compared to emerging “meta-harnesses,” like Databricks’ Omnigent and perhaps even the self-improving Exo harness.

Note: we’re also publishing our interview with Exo coauthor Alex Krentsel this weekend; it’s worth a watch and has a bonus discussion on OpenClaw architecture!

Schott rightly noted that there’s confusion about what the term meta-harness even means at this early stage. Regardless, he thinks having one API for working across all harnesses would muddle the story for Flue. His framework specifically defines how skills work in Flue, how subagents work, and so on. As he put it, “the framework [Flue] and the harness are very intertwined.”

He personally finds the meta-harness discussion fascinating, and has played with Exo, but says it’s “a different interest scenario that isn’t really related to hosted agents.”

The Cloudflare connection

Throughout the interview, Schott referenced being able to take advantage of his employer Cloudflare’s tooling and infrastructure. But he was also very clear that Flue is an “open source framework for every host,” as he put it, and he wants it to stay that way.

“The best tools are the ones that float above the host,” he said. “That opens the door for the most developer adoption and the most innovation.”

Host portability is one of Flue’s defining principles — and perhaps that’s where the fundamental difference to Vercel’s eve is. While eve can also be self-hosted, it is optimized to take advantage of Vercel’s many features. Of course that’s a known playbook of Vercel, which does the same thing with Next.js.

All that said, Vercel itself has shown that a Flue agent can be deployed on Vercel. So the two companies can play nice together.

I also mentioned LangChain’s new Managed Deep Agents offering as an example of hosted agent platforms coming onto the market. However, Schott said a managed agents product is not currently on Flue’s roadmap.

“It’s so early for us, we’re just focused on building the best harness,” he said.


Links to find Flue and Fred online; Richard is at @ricmac. This is a new written interview series we are trying out for subscribers — let us know your feedback!

  •  

Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web

One of the most watched videos from the recent AI Engineer World’s Fair is a 20-minute talk by Frank Coyle, a professor of computer science who currently teaches generative AI and LLMs at UC Berkeley. Drawing on his decades of experience, Coyle re-introduced the concept and practice of ontologies to today’s AI engineers.

Coyle argued that while LLMs are very effective at providing probabilistic reasoning, for agentic systems to be truly effective they need “logical guardrails” — by which he means ontologies.

UC Berkeley professor Frank Coyle speaking at AIEWF 2026.

In computer science, an ontology is “a description of data structure – of classes, properties, and relationships in a domain of knowledge” (as nicely defined by Oxford Semantic Technologies). Coyle himself defined an ontology as simply “data as graphs.” He added that ontologies as a concept go right back to Aristotle, and have been used throughout the history of Artificial Intelligence.

The company Neo4j, known for its graph database systems, is also using ontologies in its agentic products.

In a keynote session at AIEWF, Neo4j CEO Emil Eifrem talked about three different types of ontologies to enable a “smarter shared substrate” to run agents at scale. The first is a business-facing ontology, describing the key concepts in an organization; next is a technical ontology, which Eifrem described as “all the metadata of all the data sources and data assets in your enterprise ecosystem”; and finally execution traces, which are “the runtime signals out of your agent.”

Neo4j’s ontology-based semantic layer

What’s Old is New Again: the Semantic Web

Frank Coyle thinks web ontologies are especially useful in building agentic systems. He mentioned Schema.org, FOAF, Dublin Core and other ontologies familiar to web developers — or at least, those of a certain vintage. He also mentioned “augmenting technologies” like RDFS (Resource Description Framework Schema) and OWL (Web Ontology Language).

One benefit of using these established ontologies is that they’re already in the training data of LLMs, so developers can just prompt for them — much better than reinventing ontologies from first principles.

“This stuff has been out there underlying a lot of what we already do,” Coyle said. “So take advantage of these things that already exist.”

As an example, he showed a Claude agent loop which used an ontology to help validate the LLM’s reasoning after the tool had run.

Coyle calls the convergence of probabilistic agents with ontologies “neurosymbolic AI.”

“Sounds pretty fancy,” he said, “but it’s really neural networks tied into symbolic AI — rule-based systems come under that category, as do the knowledge graphs that we’re assembling.”

He reiterated that neurosymbolic AI represents “a way to keep the LLM on its guardrails.”

The Pros and Cons of Ontologies

Someone who has been steeped in ontologies for many years and is now combining that with AI engineering is Kingsley Idehen, who runs a company called OpenLink Software. He’s been building an “agent engineering stack” that uses Semantic Web technologies — including an “agent with RDF memory.”

Kingsley Idehen’s agent-rdf-memory system.

I asked Idehen to explain to me the benefits of ontologies.

“The beauty of LLMs is that they are powerful processors of language,” he replied. “The beauty of an ontology is that it defines the types of entities and relationships through which language acquires computable context. Together, they bring the expressive power of language to computing’s UI/UX stack.”

That makes a lot of sense, but if you’ve been a web developer for a while you’ll know the challenges of ontologies: maintenance and keeping them up-to-date. It’s why the 1990s and 2000s vision for a “Semantic Web” — which was based on ontologies — never took off.

Current AI developer Prasenjit Sarkar offered a potential solution for the maintenance problem on X, arguing that “when an agent maintains the ontology as part of its own operation, updating definitions when it encounters edge cases, the maintenance problem changes character.” It’s still a hard problem though, he added.

Despite these issues, the structured nature of ontologies does appear to be a good match with the occasionally wayward tendencies of probabilistic LLMs. You get the power of LLMs, but ontologies will keep them in check.

Plus, as Neo4j’s Eifrem explained, ontologies allow us to move from “a world of thick agents with manually wired data sources” to a new world of “thin agents on a smarter shared ontology-based semantic layer.”

Neo4j’s “thin agents” concept, which relies on ontologies.

Loops and Guardrails

Back to Coyle’s presentation. He also had a great point about loop engineering, which he noted “has been around forever” in computer science. The problem, of course, is that loops can break or otherwise “go off the rails.”

Again, this is where an ontology system can act as a guardrail to a probabilistic LLM. One of Coyle’s slides referred to it as “a bounded set of rules around an unbounded loop.”

Near the end of his presentation, Coyle demonstrated how to use OWL as a check on agents. One slide showed that while language can be slippery, “an OWL axiom is a rule a machine enforces.”

He also showed how “you can have a reasoner built on ontology, to check [and] keep the LLM on track — have guardrails to keep it honest.”

Semantic Vibes

Perhaps ontologies are starting to resonate with AI engineers because a central concern at this time is quality control for loop engineering. We saw this debate play out at AIEWF, with many conference speakers not willing to go all-in on fully automated “software factories” just yet. One of the key learnings from the event was that there need to be guardrails and humans in the loop.

Also it’s fascinating to see traditional web technologies make a resurgence in the field of AI engineering, especially after the 2025 trend of “vibe coding” made it seem like anyone could create software. Of course, since then the penny has dropped: we need to maintain that software and make sure it doesn’t break! So in 2026, we’re seeing a return to software engineering discipline — including now a revival of web ontologies as a way to keep probabilistic LLMs honest.

  •  

5 Trends That Defined AI Engineering at World’s Fair 2026

swyx’s note: thanks to Richard for covering AIE while I was working on the conference itself! Make sure you have opted into the AINews feed to get our weekday updates. AIE next returns to NYC, Oct 12-14, with a heavy focus on AI in Finance this year.


AI engineering has come a long way in three years. When swyx coined the term “AI engineer” in June 2023, he was giving a name to a new kind of developer emerging from the big bang of large language models. It seems like ancient history now, but remember when we called the intersection of AI and software development “prompt engineering”? That was just months before swyx’s reframing.

The latest AI Engineer World’s Fair showed just how much the field has matured. Whether or not “AI engineer” has become a formal job title everywhere is almost beside the point. The engineering practices that have developed around AI over the past three years — building coding agents, designing harnesses, managing context, evaluating model outputs, and orchestrating increasingly autonomous systems — are becoming part of mainstream software development.

Rather than focusing on individual announcements at AIEWF, this post will pick out five larger trends that show where AI engineering stands in 2026.

1: The focus shifts from agents to the systems around them

One of the clearest ways to see how AI engineering has evolved is to compare two essays by former OpenAI researcher, and now co-founder of Thinking Machines Lab, Lilian Weng. Her influential 2023 article, LLM Powered Autonomous Agents, described the anatomy of an LLM agent in terms of planning, memory and tool use. AutoGPT, BabyAGI and GPT-Engineer were among her examples — proof-of-concept systems that suggested autonomous agents might soon become practical.

Her new 2026 essay, Harness Engineering for Self-Improvement, takes a very different perspective. Rather than focusing on the agent itself, Weng argues that the system surrounding the model has become just as important: the harness that manages workflows, context, permissions, evaluation, persistent state and continuous improvement. In other words, AI engineering has moved beyond prompting models toward engineering reliable systems around them.

Coding agent loop; Image by Lilian Weng

This shift was very much top of mind at AIEWF. I don’t think AutoGPT — the buzzy autonomous agent project everyone was talking about in 2023 — was even mentioned this year. Instead, the conversation revolved around Claude Code, Codex, Gemini CLI, Cursor, Warp and all the infrastructure needed to make coding agents dependable in production.

I remember being turned off by the AutoGPT buzz at the 2023 event, mainly because all the discussions seemed to focus on removing humans from the equation. But over the past few years we’ve learned that complete agent autonomy is not only unreliable, it isn’t even desirable — especially at scale. So it was a relief that at AIEWF, agents were largely positioned as augmenting the AI engineer, rather than replacing them.

During the OpenAI keynote on day 2 at AIEWF, Romain Huet emphasized this point. Using tools like OpenAI’s Codex, Huet argued, engineers can more easily collaborate with agents. As he put it, “software ate the world, and then AI ate software, but now what we’re here to say is that the AI engineers are eating the world.”

Despite the growing power of AI engineers, there’s also a sense that even the frontier companies don’t fully understand how their models are evolving — and so how much control can engineers truly have over them? In a separate keynote, Anthropic’s Thariq Shihipar talked about how their latest model, Claude Fable, is like an organic system — “models are grown, not designed.” There’s a “capability overhead,” he said, where “Claude gets smarter in a spiky way.”

All the more reason to build systems for agentic development, so that we can evaluate and monitor the outputs.

2: Loop engineering is the new control layer

By the end of the first morning of keynotes at AIEWF, it was clear that “loops” was the buzzword du jour of the event. Overuse of the term aside, it did highlight a key point of tension around AI engineering: how much control should agents have, and where should humans remain in the loop?

OpenClaw creator Peter Steinberger advocating for better loops.

One approach a lot of leading engineers are now taking is putting themselves in an “outer loop” — to oversee the largely autonomous work being done by agents in an inner loop.

Roland Gavrilescu is co-founder and CEO of Introspection, a new company building infrastructure for deploying self-improving systems. In an interview with Latent Space, he explained how the concept of “autoresearch” provides the necessary feedback structure for agent loops:

“You can think of the system as having an inner loop and an outer loop. The inner loop is the primary system interacting with users and performing the work. Autoresearch is more concerned with the outer loop: another system that studies and maintains the primary system.“

The outer loop can include feedback signals, evals and human input. So it might still be largely autonomous, but the point is it is a method of oversight for the primary agent loop. Former Google engineering leader Addy Osmani had a nice line relating to this, saying that “agents can run much more of the inner execution loop, but that outer loop is still engineering.”

The term “loop engineering” came up multiple times during AIEWF, suggesting that it’s the human AI engineer’s responsibility to build these loop systems. Even the “ClawFather” Peter Steinberger, creator of OpenClaw, makes a point of putting himself in the outer loop. In the OpenAI keynote, he explained that “the agent runs the inner execution loop; I set the direction and I make decisions in the outer loop.”

The Loop Debate at AIEWF.

On the final day, an on-stage debate was held to determine whether fully autonomous agents were capable of managing loops in reality. Dex Horthy from HumanLayer claimed that “the hype is outrunning the discipline.” He wasn’t against loops, per se, noting that Kubernetes is built on control loops — “but they’re deterministic loops.” Geoffrey Huntley, creator of the Ralph Loop, admitted that loops were “frontier thinking,” but he had a wonderful analogy for the audience to ponder:

“[We’re] kind of like locomotive engineers now. That’s our job: to keep the locomotive on the rails.”

3: AI engineering enters the enterprise

This way of working with AI tools is starting to make its way into enterprises, typically via a new role called a “forward deployed engineer” (FDE) — where engineers work directly with organizations to implement AI capabilities.

Natalie Meurer, who leads FDE at Sierra, told Latent Space that implementing AI into organizations typically requires a lot of orchestration. “Every enterprise we work with wants to know how it can maintain everything its agentic ecosystem is capable of doing,” she said. “It needs to manage all the integrations and all the teams that contribute to the agent.”

Cursor’s Pauline Brunet talking about FDEs in an AIEWF session.

In her session at AIEWF, Cursor’s Pauline Brunet spoke about what their FDEs look to achieve in each engagement:

“When [we] walk away at the end of the engagements — and we, in our case, have deployed cloud agents, long-running agents, automations, [and] we’ve built applications on top of our Cursor SDK — that when we walk away, it is a strict ROI for them. That means they’re not gonna turn things off when we leave.”

Another term used regularly at the conference was “software factory.” At Cursor, “a software factory means long-running agents helping people throughout that entire process,” said Brunet. This is basically what her team of FDEs is responsible for implementing, sitting alongside their customers’ engineers.

Where human engineers fit into a software factory is a key issue for enterprises. Warp CEO Zach Lloyd explained that organizations need to choose which parts of the lifecycle to automate, and where humans should be brought into the loop.

Warp’s Zach Lloyd on building the thing that builds the product.

“You choose your repositories, the parts of the software lifecycle you want to automate, and the points where humans should be brought into the loop,” Lloyd told us, regarding his company’s new software factory platform, Oz. “Different organizations and codebases will have different preferences. Do you fully automate code review? Do you have humans review certain high-risk changes?”

Another concern for enterprises is managing their unique organizational data in AI systems. Prukalpa Sankar from Atlan spoke at the conference about “context engineering,” explaining in a tweet that it’s important to consider “​​how context flows from your business systems into a shared company brain, then out to agents, copilots, and apps through MCP, APIs, and retrieval.”

Finally, lest we think enterprises are all-in on agents, Cursor’s Brunet pointed out that enterprise adoption of AI “is still concentrated among early adopters.” So finding “the right champions inside an organization” is a challenge for FDEs at this stage.

4: Coding agents replace IDEs as the developer interface

Perhaps the biggest practical change since the first AI Engineer Summit is how developers interact with AI on a daily basis.

In 2023, AI-assisted programming largely meant GitHub Copilot completing the next few lines of code. Most developers were still writing almost everything themselves, using AI as an intelligent autocomplete. But now we have tools such as Claude Code, Codex, Gemini CLI, Cursor and Warp. These “coding agents” can typically understand a broader objective, explore a codebase, modify multiple files, run tests, debug failures and iterate on their own work before presenting it back to the developer.

In Barr Yaron’s AI engineering survey, coding agents was a key trend.

The trend of coding agents now extends to web development too — with the recent release of Vercel’s eve, which the company calls an “agent framework,” comparable to its popular open source React framework, Next.js.

Vercel’s Chief of Software, Andrew Qu, told Latent Space at AIEWF that agents are effectively a new type of software. “They [agents] are not as predictable as web applications,” he explained. “The infrastructure can look similar, but the interaction, interface and outputs are much more dynamic.”

Qu added that the job of building a framework for agent development is far from over. “A year ago, we did not know sandboxes would become so important, or how much demand there would be for secure code execution and long-running jobs,” he said. “As we learn more from production, there will be much more to build.”

A for agents? Andrew Qu flashes the Vercel triangle logo.

This brings us back to the software factory trend, when developers are managing multiple agents. Charlie Holtz, CEO of Conductor, reminded the AIEWF audience that regardless of the coding harness, human engineers should always remain in control.

“I don’t want the future to be built around factories,” Holtz said. “I want to feel like a human, I want to be in the flow, I want to be in front of an orchestra, waving my baton.”

There was a sense during the conference that AI engineers aren’t yet aligned on which term is more appropriate: software factories or orchestras? Even Geoffrey Huntley, a loopmaxxing advocate, cautions about getting ahead of ourselves when it comes to automation:

“My biggest concern is that this time next year at the conference, we’re going to see a whole bunch of folks saying, our factories failed, our loops failed. These are things that we are still yet to figure out.”

5: Every agent platform is building around skills

One of the talking points of the conference was “skills,” a concept Anthropic popularized when it introduced “agent skills” to Claude last October. To borrow Addy Osmani’s definition, skills “encode the workflows, quality gates, and best practices that senior engineers use when building software.”

At AIEWF, Vercel’s Andrew Qu said that skills were “useful as portable, on-demand knowledge.” Introspection co-founder Roland Gavrilescu declared that AI engineering has shifted “from agent tools to agent skills.”

In a session on the main stage, Philipp Schmid from Google DeepMind showed how using skills (and other declarative Markdown files) allows developers to use “agents without code.” His main point was that skills reduce the need for orchestration code, which up till recently was typically done using Python. His conclusion:

“Agents are just files. We write markdown files to extend capabilities. Agents can learn from those, can create their own files.”

Paul Bakaus, who used to work for Google but now runs a company called Renaissance Geek, has created an entire project around agent skills. Impeccable is an open source design skills system that gives coding agents a vocabulary for improving interfaces. He even advocates for “skill engineering” as a discipline in its own right.

Paul Bakaus: “You can’t one-shot design.”

In an interview with Latent Space, Bakaus argued that most skills — and indeed most models — are not very creative. “They converge in one direction, and if everybody uses the same skill to do frontend design work or something like that, everything ends up looking the same,” he said.

Apparently there’s also such a thing as “skills hell,” which Matt Pocock said is comparable to previous developer frustrations — like frameworks hell. In a virtual presentation, Pocock provided a detailed checklist for writing skills, which you can see in the video below. In a nutshell, he advises writing fewer and smaller skills, and putting more thought into structure.

In a closing keynote, Y Combinator president Garry Tan implored the audience to use skills and other “AI native” approaches at their own startups or employers. Talking about business functions like sales, support and finance, Tan said that “the AI native companies that I see inside YC encode all of that as skills, written procedures that their agents execute, and they hire engineers whose job it is to maintain those skills, to do the work the skills can’t do yet.”

But again, there’s a danger in relying too much on what agents autonomously do. As AIEWF attendee Tyler Brown noted on X, “autonomy without structure creates as much slop as leverage.” One of his learnings from the conference was to “re-visit and re-implement your skills”:

“Each time there’s a new model release, it’s as if you have a kid that grows from middle school to high school. You have to change the curriculum for them to get the benefits of the new model.”

Agent engineering at scale

It’s been three full years since The Rise of the AI Engineer and the first AI Engineer Summit. Looking back, it really is striking how much the conversation has evolved. Three years ago, the focus was on proving that LLMs could act as autonomous agents at all (and the answer at that time was usually no). AutoGPT, prompt engineering, and early orchestration frameworks like Langchain dominated the discussion back then.

Now that agents not only work, but have proven they can scale, this year’s AI Engineer World’s Fair was able to concentrate on the bigger problems: building reliable systems, orchestrating teams of agents, managing context, evaluating outputs and integrating AI into production software.

Agents are everywhere now…even on the back of San Francisco buses.

The term “AI engineer” may have started life as a new job title, but at AIEWF 2026 it felt more like a description of where software engineering itself is heading. Whether developers call themselves AI engineers, software engineers or Forward Deployed Engineers, they’re increasingly working with the same set of ideas: coding agents, harness engineering, designing loops, and orchestration.

  •  

AIEWF Daily Dispatch: The great loops debate and the state of AI engineering

One of the highlights of the final day of the AI Engineer World’s Fair was a debate about loops. It nicely captured an argument running through the whole conference: are autonomous software factories viable now, or is the engineering discipline lagging behind the ambition?

Allie Howe from Keycard was the moderator and she opened by asking, “is there or is there not a delta between the hype behind loops and what actually works in practice?”

The pro-loop case was presented by Geoffrey Huntley, creator of the Ralph Loop, and Keycard CEO Ian Livingstone. Huntley opened by saying loops are already here. “It’s inevitable, it’s here to stay,” adding that “I don’t see myself going back to writing code by hand.”

Livingstone said that verifiability is ultimately what it’s about — and you can achieve that with any code, regardless of how it was produced. He also pointed out that loops have always been a core aspect of software development:

“A loop is at the core of ‘I try something, I learn something, I apply something.’ And all we’re really talking about is how quickly we can expedite that process.”

On the skeptical side were Dex Horthy from HumanLayer and Greg Pstrucha from Subroutine. Horthy began by noting that he wasn’t anti-loops. “The basic take here is not whether loops are good or bad,” he said, noting that “Kubernetes is actually built on loops — built on control loops. But they’re deterministic loops.” Horthy’s issue is that “the hype is outrunning the discipline.”

“I haven’t seen proof that we are at a point where we can just step up an abstraction level,” Horthy said, referring to agents controlling the coding. “I actually think we need to step down an abstraction level, if anything.”

Pstrucha was mainly concerned about the economic viability of agentic loops, which he said wasn’t sustainable. You can’t “orchestrate your problems away by buying more tokens,” he said.

“[We’re] kind of like locomotive engineers now. That’s our job: to keep the locomotive on the rails.”
- Geoffrey Huntley, loops advocate

Huntley then offered this wonderful analogy for loopmaxxing: “[We’re] kind of like locomotive engineers now. That’s our job: to keep the locomotive on the rails.”

The discussion turned to software factories, the metaphor that has really taken hold of the industry. Horthy worries that when everything is automated in a factory-like agent environment, “you never touch the problem.” So instead, he advises to start small and iterate with agent loops — to “build up intuition” and not try to automate end to end from the start.

Even Huntley recognized some of the dangers in loops. He said that software factories represent where we are headed in the future, but cautioned that it’s not yet solved in the market. “This is frontier thinking,” he said.

At the end of the hour-long debate, Howe polled the audience to ask which side ‘won’. Ironically, this resulted in a human failure: the stage lights were too bright for Howe or any of the debate participants to see how many hands were raised. If only an agent was in charge of dimming the lights.

Anthropic’s next big thing: Claude Tag

Perhaps one example of a company moving to a software factory model is Anthropic. Mike Krieger, one of the co-founders of Instagram back in Web 2.0 and now Head of Labs at Anthropic, was interviewed by swyx in one of the morning sessions.

Krieger talked about Claude Tag, Anthropic’s internal model which the company announced to the world last week. He described Tag as more delegated, asynchronous and proactive than Claude. It perhaps suggests what an early software factory looks like in practice — not agents replacing a team, but multiple people delegating responsibilities to a system like Claude Tag.

Mike Krieger talking with swyx at AIEWF today.

“Most usage is actually much more delegated,” he said regarding his team’s usage of Tag. He gave an example of how they instruct the agents: “Don’t just fix this bug. Now you are responsible for this part of the codebase, and I want you to monitor this feedback channel and proactively take on tasks.”

“That’s really changed how we operate currently,” he continued. “It’s much more this multiplayer, async, proactive way.”

However, he also indicated there are some negative consequences to becoming more automated. He noted that his team is “bottlenecked on reviews” and on the “human ability to fully conceptualize what we’re doing.”

2026 AI Engineer Survey

Back to the current reality for most AI engineers. This morning, Barr Yaron from Amplify presented her annual survey of the industry.

According to Amplify’s data, 95% of respondents now use agents — roughly double last year’s share. Among teams using agents, 89% said those agents could write data, up from 52% the previous year.

“Agents are no longer reading, summarizing, drafting,” Yaron said. “They’re taking actions inside the systems.”

Barr Yaron presenting her AI engineering survey.

The controls, however, remain comparatively primitive. Human approvals and permissions were the two leading safeguards, followed by a scattered collection of task decomposition, retrieval, memory and sandboxing techniques.

“Nobody has settled the control layer for agents,” Yaron said.

Cost is also a concern. Forty percent of respondents said that AI costs regularly limit how ambitiously they use AI, while another 36% said it sometimes does. Token usage is now the second-most monitored production metric, behind quality.

The survey captured the conference’s central contradiction. AI has made experimentation cheaper and enabled teams to produce more software, but 59% of respondents to the Amplify survey fear that today’s AI-generated code is creating long-term liabilities.

Closing keynotes

The final sessions of the conference appropriately took us back to thinking optimistically about AI technology — about building with it. After all, that’s why the AI Engineer World’s Fair exists, and it’s where the fun is!

Theo Browne showcased several software projects he had built, or was still building, with AI. His point was that the scale of what an individual developer can realistically attempt has shifted. “What used to be a startup is now a side project,” he said, while projects he would once have dismissed as “too big” are moving within reach.

Garry Tan, president and CEO of Y Combinator, followed by giving that optimism an organizational form. The fastest-growing founders YC sees, he said, are “not treating AI as autocomplete, they’re treating it as a workforce.”

Garry Tan at AIEWF.

Tan’s closing prescription was: “Build an AI-native company, not a company that just uses AI.”

The debates during the week showed how much engineering remains before the AI-native vision is viable for all. But the closing keynotes offered a reminder of why the engineers who attended this conference are pursuing it: they just want to ride those locomotives!

  •  

Vercel's Andrew Qu on why agents are a new kind of software

Vercel’s Andrew Qu on the AIEWF expo floor.

Andrew Qu is Chief of Software at Vercel, where he works with the CTO across internal engineering, product experimentation and emerging technologies. He has built libraries for MCP, created skills.sh and led the development of eve, Vercel’s framework for building agents.

In this interview with Latent Space, Qu explains why agents represent a new form of software, what Vercel learned from building its own, and why Vercel itself is turning into an agent!

From web applications to agents

Latent Space: What does a Chief of Software do at Vercel?

Andrew Qu: My role is pretty unique. I work with the CTO to ship impact in any way, shape or form. It’s a mix of internal engineering, external experimentation and staying on the frontier by building things.

That means building new libraries and frameworks and showing people how to do things for the first time. I built an MCP library that made it easier to create some of the first MCP servers, and I also built skills.sh to make agent skills easier to discover and use.

Latent Space: How did Vercel evolve from focusing on web development to investing heavily in agents?

Qu: Vercel’s origins were about making it easy for developers to ship websites and web applications. More recently, we’ve seen a shift from people building pages to people building agents.

While building our own agent in v0, our vibe-coding product, we ran into a lot of paper cuts that existing tooling did not solve: switching models or providers, adding fallbacks and making runs resumable.

We turned those solutions into reusable libraries that could support v0 and also help customers build their own agents. Over time, we accumulated a set of primitives and decided to assemble them more cohesively. That became eve.

Why eve became necessary

Latent Space: How did you reach the point where Vercel needed a dedicated agent framework?

Qu: About a year ago, I started working toward putting an agent on every desk inside Vercel. That led me to build a successful data agent, and along the way a number of best practices emerged: filesystem agents, skills, compaction and subagents.

These were all things I wished had come out of the box. Eventually, we asked: what if there were a prescriptive way to do this, so other developers did not have to go through the same exploration? That is where eve came from.

Latent Space: Are agents simply another kind of application, or a genuinely new form of software?

Qu: I think agents are a new type of software. They are not as predictable as web applications. The infrastructure can look similar, but the interaction, interface and outputs are much more dynamic.

That changes how you build them. You need different primitives for context, tools, resumability and long-running work.

Latent Space: What kinds of problems are particularly well suited to agents?

Qu: We see a lot of business agents. Internally at Vercel, we use them for repetitive work ranging from a first pass at legal contract redlining, to marketing retrospectives and identifying people to contact, to writing queries against our data stores.

A good candidate is often a repetitive task that still requires some reasoning. It is not just fixed automation, because the system has to interpret the situation and decide what to do.

Building effective agents

Latent Space: When should an agent work autonomously, and when should a human remain in the loop?

Qu: I don’t think the future is all autonomous loops, and I don’t think it is all human-in-the-loop. It is about choosing a feedback cycle that fits the task.

If the task is well defined and you know what the final output should look like, it can be reasonable to let a loop continue until it is done. For more careful or surgical engineering work, you should check back in and make sure you are steering the model correctly.

Latent Space: Your approach evolved through prompting, bespoke tools, coding-agent harnesses, filesystem agents and skills. What was the main lesson?

Qu: We are still figuring out what makes an agent productive. Along the way, we have been collecting these primitives and bringing them together in eve.

There will be more to add as best practices emerge. A year ago, we did not know sandboxes would become so important, or how much demand there would be for secure code execution and long-running jobs. As we learn more from production, there will be much more to build.

Latent Space: Is Vercel creating an end-to-end agent platform comparable to the one it built for web development?

Qu: Yes and no. We value partners that provide specialized parts of the agent lifecycle, but we also want it to be very easy for developers to get started.

If you deploy eve to Vercel, you get observability and evaluations out of the box. We want to make that experience more comprehensive while making it easy to integrate with partners rather than owning every component.

Skills and current knowledge

Latent Space: Why have skills become so important?

Qu: Skills are useful as portable, on-demand knowledge. Models often contain outdated information. For example, they still sometimes recommend Vercel Postgres, even though we deprecated it years ago in favor of our marketplace.

A skill can tell the agent that Vercel Postgres is deprecated and steer it toward the current approach. Until companies can audit and update every old piece of content, skills provide a way to forward-correct the model.

I would recommend publishing skills for the latest version of your product. But companies should also audit their existing content, identify what is outdated and update it or add clear notes.

An agent-readable web

Latent Space: How will websites evolve as more traffic comes from agents?

Qu: We have published reports showing bot traffic rising while human traffic is stagnant or declining, even as impressions increase, because agents and bots are hitting websites more frequently.

The future of the web is therefore to be as accessible to bots and agents as possible, so they can learn about your product and use it successfully.

At Vercel, we already detect when an agent makes a request and serve Markdown directly. Instead of forcing it to process HTML designed for a visual browser, we provide a format that is easier to read.

Latent Space: Does that mean one experience for humans and another for agents?

Qu: I think so. Humans may continue to receive the visual site, while agents receive a more structured, machine-readable representation. We are already doing that today.

What comes next

Latent Space: What problems are you most interested in solving next?

Qu: One of the things at the top of my agenda is multiplayer agent development. Whenever a team collaborates, people struggle to share context.

I may have techniques for getting a front-end interface right on the first attempt, but another person may not know them. I am interested in how we can share that context between teammates and allow them to contribute to it.

Latent Space: Will agents become a separate application category, or a standard capability built into most software?

Qu: It depends on who you are and what you are building. For Vercel, Vercel itself is becoming an agent. We have an agent on the website, in Slack and in the dashboard that can do things on your behalf.

Other companies will ship agents as standalone products. For us, agents are tightly coupled to everything we build. We want the entire platform to be agent-friendly — and, in many ways, to make the platform itself an agent.

  •  

The website of the future may assemble itself for every visitor

Adobe Principal Scientist Carlos Sanchez at AIEWF.

For as long as I can remember (and I managed websites in the dot-com period), “personalization” has been a holy grail for websites. But up till now, that’s typically meant selecting from a predefined set of options. A retailer might recommend an item based on a previous purchase, or place a visitor into one of several audience segments — that’s been the extent of personalization.

Adobe Principal Scientist Carlos Sanchez is exploring a more radical possibility: what if the website itself could be assembled around the needs of each visitor?

At the AI Engineer World’s Fair in San Francisco, Sanchez demonstrated what Adobe calls an “agentic site” — a web experience that interprets a visitor’s intent, retrieves relevant material from the company’s existing content, and composes a personalized page in real time.

Adobe calls this approach an “audience of one.” Sanchez’s larger point was that the technology is no longer hypothetical.

“Many people don’t even think it’s possible to generate a web page on the fly,” he told Latent Space after his session. “People think it is future-looking. No, you can do this. It’s not the future, it’s the present now.”

From personalized components to personalized pages

During his presentation, Sanchez demonstrated a site that used the visitor’s browsing behavior and search queries as signals. The system grouped those signals into an intent category — such as exploring, researching or preparing to purchase — and then used an LLM to assemble a page suited to that intent.

In one example, a visitor interested in camping received a version of a coffee-machine site whose copy, product selection and supporting content had been reorganized around making coffee outdoors.

Sanchez also showed a more open-ended interface in which someone could enter a query such as “Europe AI conferences” and receive a page composed specifically around that request.

“We call this ‘audience of one,’ because the idea is to personalize the site in real time based on the user accessing it and what the user is doing,” Sanchez said.

The idea is that the site’s existing content is the grounding corpus. Adobe’s system retrieves from that material rather than asking an LLM model to invent an entire experience from scratch.

For AI engineers, one potential constraint is latency. In his session, Sanchez said that Adobe evaluates models not only for accuracy, but also for speed: “We don’t want the site generation to take more than one or two seconds.”

Sanchez says the economics are already becoming plausible. He estimated the current inference cost at “one to two cents per page.”

“But our point is also this is only going to get cheaper,” he said. “This is where we are today. In six months, who knows where we’re going to be.”

AI makes it easier to build, but harder to choose

Adobe has not yet broadly deployed these experiences on production customer sites. Sanchez said the company is presenting the concept to customers and looking for organizations willing to experiment.

Commerce is an obvious initial use case, because personalization can be connected directly to conversion. But the opportunity is not necessarily limited to retail. “It could work for other things — anything that needs more conversion and has a big matrix of user types or personas,” he told me.

Still, Sanchez acknowledged that he’s unsure if agentic sites will become a widespread reality.

“With AI, it’s very easy to build things, but it’s hard to know what to build,” he said. “We build things and then we find the customers.”

It’s not just Adobe feeling the uncertainty around its ‘audience of one’ concept. Website owners are currently evaluating all kinds of AI functionality: chat interfaces, structured content (like WebMCP), generative UI, personal agents, and more. Not to mention trying to find ways to bring users in from third-party AI platforms.

“I think it’s a combination of all these crazy different ways,” Sanchez said. “You are in a chat, I want to show UI, I want to get you to buy something. Then you’re in a site, I want to steer you this other way. Maybe you’re in an OpenAI chat and I want to bring you into my site. Everybody’s trying to figure this out on the marketing side.”

A web built for humans — and agents

Of course, websites in 2026 and beyond won’t just be personalized for human visitors.

As personal agents become more capable, a user may delegate some purchases or research tasks entirely. The agent could arrive carrying a much richer expression of the user’s preferences than the destination site could infer from cookies or recent browsing behavior.

Sanchez expects websites to evolve for both kinds of visitor. “Whether it’s going to be two versions [of a website] or not, that may be blurry,” he said. “But obviously, you’re going to have to target both.”

Also, not every transaction will work the same way. A personal agent might autonomously reorder toilet paper, while a person buying a jacket may still want to inspect the product and make the final choice through a visual interface.

That means websites will need to support different levels of delegation and involvement, rather than treating “agentic commerce” as a single interaction pattern.

Technologies such as WebMCP could allow a site to expose structured tools directly to an agent, while MCP Apps and other generative interfaces could bring interactive product experiences into the user’s chat environment. An A2A backend might allow agents to interact without traversing the conventional visual site at all.

It might end up being one site with both visual components and agent-accessible tools — two distinct experiences — or perhaps a human-facing website paired with an agent-to-agent service.

“That’s still what everybody’s trying to figure out,” Sanchez said. “But there’s going to be agentic targeting, for sure.”

Whither websites?

Whether websites survive the AI era at all is another big question we’re all grappling with.

What I gleaned from Sanchez at AIEWF was that the traditional website is unlikely to disappear completely, but its role will surely change.

Rather than being a fixed collection of pages that every visitor navigates, a “website” could become a governed content and interaction system that assembles an appropriate interface on demand. At least, that’s the future that Adobe is actively exploring.

  •  

Skill engineering and the case against one-shot AI design

Impeccable’s Paul Bakaus at the AI Engineer World’s Fair.

Paul Bakaus thinks the emerging discipline of “skill engineering” can make AI agents more capable — but he absolutely does not want to remove people from the creative process. He chats to Latent Space about his approach to design in the AI age.

Bakaus is the creator of Impeccable, an open-source design skills system that gives coding agents a vocabulary for improving interfaces. Instead of asking an agent to redesign an entire website in one shot, users can tell it to make a section “bolder,” “quieter,” “denser,” or more polished.

Behind those apparently simple commands is a larger argument about how AI products should be built. Agents need more than instructions, Bakaus said: they need domain knowledge, context and carefully defined ways for humans to steer the result.

“The point is to give you a way to steer what you want to end up with,” he said during a session at the AI Engineer World’s Fair. “It’s never going to be a tool for one-shot design. That’s not the intent.”

The emerging craft of skill engineering

Impeccable began as a relatively simple extension of Anthropic’s frontend design skill. As its audience grew, Bakaus expanded it into a more complex system with multiple components and workflows.

That process led him to start thinking of skill engineering as a discipline in its own right. His workshop at the conference explored what he called the “dark arts” of building skills.

“One of the interesting topics was that most skills — [and] most models — are not very creative,” Bakaus told me. “They converge in one direction, and if everybody uses the same skill to do frontend design work or something like that, everything ends up looking the same.”

Skill engineers must also account for differences between agent harnesses and models. Codex and Claude, for example, do not necessarily handle subagents or permissions in the same way. A skill intended to run across Claude Code, Cursor, GitHub Copilot and Codex cannot assume they all provide identical capabilities.

Bakaus has also experimented with routing inside a skill, allowing it to combine several capabilities and direct a task toward the relevant instructions. He compared this to a mixture-of-experts model, with routing used both to conserve tokens and improve effectiveness.

Giving agents a design vocabulary

Impeccable’s core innovation is to take terms familiar to designers and give them a more precise operational meaning for an agent.

An unassisted model asked to make a page “bolder” may add gradients, neon effects or glass-like surfaces. Impeccable instead defines boldness through concepts such as hierarchy, scale and decisive typography — changes that attract attention without necessarily breaking the existing design system.

“An adjective with nothing behind it is just a nice apostrophe,” Bakaus said. “You really have to tell the agent what you mean.”

He described these terms as words that have been “imbued with meaning.” The model already has some conception of what words such as “bold” or “quiet” mean, but the skill translates them into a specific professional domain.

This is the key, because experts often possess a vocabulary that non-experts do not. Bakaus said he had observed large differences between the work produced by a designer and an engineer using the same model, simply because the designer knew how to articulate the desired result.

“I’ve been trying to put that language — basically compress it into a skill and into a system — to be able to express yourselves better,” he said.

However, he does not believe every part of design can be controlled from this level of abstraction. Directly manipulating spacing may still be the fastest option for a small adjustment, while open-ended prompting can be useful during initial exploration.

The objective is not to replace every tool with an agent, he insisted. It is to determine “the exact level of control” and insert the person at the point where their judgment is most valuable.

Designers and engineers move up the stack

Bakaus sees the boundaries between design, engineering and product management becoming less distinct.

“Designers are moving into code, engineers are moving into design, and vice versa,” he said. “These worlds are all colliding.”

That shift will be uncomfortable for people whose work primarily consists of translating an existing artifact into another form. Engineers who mainly turn Figma designs into code face growing automation, while designers whose contribution is limited to making an existing interface look competent face similar pressure.

“Designers all have to move one layer up the stack to think more about the what,” he said. “I think the role of the product manager and designer is actually converging.”

At the same time, designers are moving closer to implementation — into code. Bakaus initially expected Impeccable to appeal mostly to engineers and assumed professional designers might resent that. Instead, he estimates that designers now make up at least half of its audience.

“So rather than moving directly into code and, you know, having no help,” Bakaus said about designers, “they use Impeccable as a bridge, because it communicates the way they communicate. And that was not obvious to me when I first built it.”

Impeccable also has a live mode that combines visual selection with an underlying coding agent. A user can select a section inside a development environment and request several alternative layouts or (for example) ask for a bolder or quieter treatment. The system operates within the project’s existing code and design system rather than exporting an isolated mockup from a third-party design tool.

Bakaus described this as a potential “design harness” at the intersection of chat and direct visual manipulation.

There will be no auto mode

The AI industry often treats complete automation as the natural endpoint of product development. Bakaus rejects that premise.

He sees two dominant camps: people trying to preserve the traditional Figma-centered workflow, and on the other side advocates of “loopmaxxing” who want agents to work with as little human intervention as possible.

“The truth is somewhere in the middle,” he said.

His preferred model is for AI to produce the first 80% quickly: the competent layout and basic implementation that would otherwise consume a lot of time. The person then owns the final 20%, where taste, context and a distinctive point of view enter the product. This is a key part of Bakaus’s design philosophy in the agentic era.

“People need purpose, and they want to play a role in whatever they create,” Bakaus said. “When you work with the agent, then you feel more ownership of the product.”

Users regularly ask him to add an automatic mode to Impeccable so that the system chooses the commands itself. He has no intention of doing so.

“There is no auto,” he said, “and there will be no auto.”

Asked about the language of software factories and other visions that appear to remove people from engineering altogether, his response was unambiguous.

“I’m squarely against that.”

  •  

AIEWF Daily Dispatch: Autoresearch and the tension between AI and human agency

“You can’t one-shot design.” Paul Bakaus at AIEWF today.

Wednesday was autoresearch day on the AI Engineer World’s Fair main stage.

Autoresearch is — you guessed it — a kind of loop. Introspection co-founder Roland Gavrilescu explained it best in an interview with Latent Space this morning. He said autoresearch “allows you to build loops in which agents help maintain the system itself.” He called it an “outer loop” that “studies and maintains” the primary, inner loop.

While autoresearch was not specifically mentioned by Anthropic’s Thariq Shihipar, who works on Claude Code, his keynote reflected the same idea of continuous discovery and adaptation. “The models are grown, not developed,” he said. “We sort of figure out and learn with the model as we use it.”

Anthropic’s Thariq Shihipar at AIEWF.

Former Google engineering leader Addy Osmani also spoke about loops, but his framing differed sharply from Gavrilescu’s.

Where autoresearch puts agents into the loop that studies and maintains the system, Osmani argued that the outer loop should remain human. “Agents can run much more of the inner execution loop,” he said. “But that outer loop is still engineering.” His summary was even more direct: “That inner loop is capability. The outer loop is agency.”

Addy Osmani’s Agency Ladder

Human agency is still important

This tension between what agents should do and what human engineers should retain was a recurring theme throughout the day. I also detected some pushback against the “software factory” framing that dominated Tuesday. This tweet from Notion’s Geoffrey Litt summed it up:

Litt drew a large audience in the Design Engineering track today, where he spoke about “how and why humans need to understand our code.” Lily Zhang tweeted the key takeaway: “The future will be very polarized: those who understand will keep having the next big idea. Those who delegate understanding will be replaced by the agent.”

Later, Litt posted a thread expanding on his argument. Although he acknowledged that agents are increasingly capable of handling more of the process, humans still need to understand what is happening. “You can learn what the agent is doing to make sure you can be an active participant in the creative process,” he wrote.

Another AIEWF speaker seeking to reinforce human agency was Paul Bakaus, who ran a session about his new design tool, Impeccable. Bakaus rejected both extremes: continuing to design entirely by hand, or “loop-maxing” toward a fully hands-off process. “The truth is somewhere in the middle,” he told me after his session.

His goal is to let agents handle the laborious first 80% of the work, before bringing the human back in “for the last 20% to make it a unique thing — to really put in your taste, your point of view.”

“There is no auto, and there will be no auto.”
- Paul Bakaus, Impeccable

For Bakaus, that is not simply a temporary limitation of today’s models. It is also about authorship and accepting responsibility for your work. “People need purpose, and they want to play a role in whatever they create,” he said. “When you work with the agent, then you feel more ownership of the product.”

This philosophy is built into Impeccable itself. “There is no auto, and there will be no auto,” Bakaus told the audience. What he means is that his product will never “one-shot” a solution — the user must be involved in the design process. “The point is to give you a way to steer what you want to end up with,” he added.

Generative media

The same question surfaced during a panel on generative media. As image, video and audio models become more capable, the issue is not merely what they can generate, but whose judgment shapes the result.

Nicole Brichtova, who works on Google’s generative media products, including Nano Banana, drew a distinction between average preference and cultivated expertise. “Somebody who has honed a craft has a very different level of expertise,” she said. “You see things that the average human will not.”

This matters because every model has a default aesthetic, whether its creators acknowledge it or not. “It ends up being us,” Brichtova said. “It ends up being the modeling teams.” She suggested that model developers may need to work more closely with people who have “a really creative point of view” — effectively bringing the art director back into the loop.

Shane Gu made the same point more broadly. Even as models become better at generating and refining their own outputs, he argued, humans must retain the sensitivity to notice what is wrong, generic or insufficient.

“Maybe right now the AI can do a lot of all the promptings and it’s sufficient, but if it’s like that, never be satisfied [that] AI is generating the content. Always find your sensitivity.”

Agentic sites

Even the web itself — the ultimate human information network — is grappling with how much automation to use.

In his session this afternoon on “agentic sites,” Adobe principal scientist Carlos Sanchez demonstrated websites that assemble and personalize pages in real time based on a visitor’s intent. He presented this transition as increasingly inevitable: “This is now possible. It’s only going to get better. It’s only going to get cheaper. It’s only going to get faster.”

But Sanchez also sounded a note of caution. “With AI, it’s very easy to build things, but it’s hard to know what to build,” he told me afterwards. That becomes especially important when an agent is generating experiences on behalf of a brand. “You cannot just generate the whole site,” he said, because the result may stray outside the brand’s guidelines.

That brings the discussion back to autoresearch. Agents may increasingly be able to observe, evaluate and improve other agents, but humans must still define the goals, judge the results, and take responsibility for what the loop produces.

As impressive as agentic technology is now, and as compelling an idea as automated “software factories” might be, you still need humans in the loop.

  •  

Autoresearch: The feedback loop behind self-improving agents

Introspection’s Roland Gavrilescu at AIEWF.

We’ve heard a lot about loops at the AI Engineer World’s Fair this week. Another buzzword is autoresearch, which involves building an “outer loop” where agents help maintain and improve the primary system, using feedback signals, evals and human input to make progress over time.

At least, that was the framing of Roland Gavrilescu, co-founder and CEO of Introspection — a new company building infrastructure for deploying these self-improving systems. Before starting the company, Gavrilescu worked on agent infrastructure and cloud agents at xAI, where he met his co-founder, Julian Bright.

Ahead of his “Autoresearch in the Wild” session at the AI Engineer World’s Fair today, I spoke with Gavrilescu about the shift from agent harnesses to feedback loops, the role of the open-source Pi framework, and why autonomous software factories must first learn from humans.

From xAI to Introspection

Latent Space: How did your new company, Introspection, come about?

Roland Gavrilescu: Last year, I was at xAI, where I met my co-founder. We were working on agent infrastructure and cloud agents, and we felt there was a new agent form factor that needed to be explored further. xAI was not necessarily the environment where we could focus completely on that.

We decided to leave and ask what a company designed around this new form factor might look like. We were interested in what made companies such as Cursor and Cognition successful, and how we could turn some of those ideas into a product that others could use.

That became the basis for Introspection.

Autoresearch allows you to build loops in which agents help maintain the system itself. The challenge is designing the right signals and feedback mechanisms so agents can improve the system, make architectural decisions and move in the right direction without constantly being bottlenecked by humans.

The loop becomes the product

Latent Space: Your session is titled “Autoresearch in the Wild” — what will it cover?

Gavrilescu: We have heard a lot about what autoresearch can do for improving experiments, but we wanted to talk about what these loops look like in production.

We are presenting three patterns that we think form the basis of a new blueprint.

The first is that the loop is the product. We have moved from focusing on models, to harnesses, and now to loops. The key question is whether you can define the right feedback mechanisms so agents can take on more work without generating more slop.

The second pattern concerns what the loop generates and how you track it over time. We are proposing a concept called an agent recipe.

We moved from agent tools to agent skills. Recipes are a larger container that brings together the components needed to encode human expertise: evals, judges, signal processing and the information that feeds back into the loop.

The goal is to create a portable format that agents can iterate on, almost like a research laboratory, but in a provider-agnostic way.

The third pattern is about what we optimize for. How can the system become both better and cheaper over time?

Companies such as Cursor and Cognition have shown that these products can work. The next stage is making them more accessible, faster and cheaper, and gradually distilling the capabilities of frontier models into systems that you own and that are customized for your environment.

Agent recipes

Latent Space: Can you explain more about what an agent recipe is…

Gavrilescu: It’s like a description of the ingredients you need and how they evolve.

The idea comes partly from data recipes used in model post-training. A data recipe describes how much data from different domains should be baked into a model.

Agent recipes are similar. A recipe might describe how your harness works with different models, the evals you use, the judges you have created, the human expertise you have captured and the failures that led to new evals.

Imagine that tomorrow you suddenly gained access to the Devin codebase. The code alone would not necessarily be that helpful if you could not see how the team arrived at the current version. You would want to understand the failures, mistakes and decisions that informed it.

A recipe captures that process. You begin with a baseline and then record how each signal produced a new judge, embedded new human expertise or led you to introduce a different model.

The inner loop and the outer loop

Latent Space: Does autoresearch mean orchestrating multiple agents, or can it involve one agent repeatedly working and verifying its results?

Gavrilescu: You can think of the system as having an inner loop and an outer loop.

The inner loop is the primary system interacting with users and performing the work. Autoresearch is more concerned with the outer loop: another system that studies and maintains the primary system.

The question is how to design that outer loop so it makes progress on the right problems without consuming an unreasonable number of tokens while deciding what to do.

Pi as the Linux of agent harnesses

Latent Space: You have compared Pi to Linux. In that analogy, is Introspection something like Red Hat?

Gavrilescu: Pi is like the Linux of agent harnesses. Linux has distributions such as Ubuntu, but the underlying system is designed to be extended. Pi is similar: it was never intended to be run as an unchanged, vanilla product. Pi separates the agent loop from its extensions and configuration, which makes the agent portable. You can spin up several different agents by loading different files into the runtime.

We saw an opportunity to combine that extensibility with recipes and open-source building blocks that can evolve for each customer while remaining portable and easy to deploy.

Making loops reliable in production

Latent Space: Reliability and the messy reality of agent loops have been recurring themes at the conference. How does Introspection address those problems?

Gavrilescu: The product is designed around the point at which you are ready to move into production.

You need to know what infrastructure is required to make the loops work, keep costs under control and maintain security. The managed infrastructure covers what is necessary for these systems to operate in production.

A major part of our focus is bringing the kind of infrastructure available inside frontier AI laboratories to a product that other companies can deploy.

Humans remain part of the system

Latent Space: What about the human in the loop?

Gavrilescu: These loops are designed with humans in the loop because you need the right signals as the system makes progress.

The human can effectively become a tool and a source of signals. Agents can be trained to ask people questions through an “ask a human” tool.

During its first few loops, an agent may rely heavily on asking questions and learning what a human would do. Over time, it accumulates those preferences and can become increasingly autonomous.

It is similar to an employee joining a new company. Initially, that employee asks a lot of questions. As they learn how the organization works, they can make more decisions independently.

Taking agent infrastructure into vertical markets

Latent Space: So what kinds of use cases are you seeing?

Gavrilescu: We are concentrating on vertical agents.

Coding agents are clearly working, and we have seen a number of companies succeed in that area. The next question is how to deploy agents in vertical and non-coding domains.

Companies in those markets are asking how they can do this securely without becoming dependent on a single provider. They want the deployment to belong to them, they want to retain ownership of their data, and they do not want to be locked into OpenAI or Anthropic. Introspection is intended to provide infrastructure that addresses those requirements using open-source building blocks.

Frontier AI labs have developed sophisticated internal agent technology. We want to bring similar capabilities into vertical SaaS and services businesses.

Why the work happens in Git

Latent Space: Is Introspection mainly intended for developers, or will product managers and other business users work with it?

Gavrilescu: We are initially focusing on software engineers in vertical SaaS companies.

We want the environment to be agent-friendly, meaning agents can work inside their own repositories and codebases. Everything is Git-based, and Git becomes the audit log that you maintain over time.

In the future, there will be interfaces that enable product managers and others to participate. But we are already seeing product managers move closer to code.

We think the right initial form factor is a human-to-agent interface in which the actual work and its history live in Git.

From orchestras to software factories

Latent Space: Does Introspection fit within the broader idea of software factories?

Gavrilescu: Yes. Designing the loops is essentially designing the factory. The remaining question is how much autonomy the factory should have.

There has also been discussion about “orchestras, not factories.” That distinction is really about the level of autonomy.

An orchestra might retain a human conductor who controls how the loops operate. A factory implies something more fully autonomous.

But you should build toward the factory rather than assume you can create a completely autonomous factory on the first day. Models do not initially possess all the context or understand every decision people inside an organization make. You cannot simply capture all of that knowledge in a Markdown file.

The right approach is to design the human as a core component of the factory. The early system should extract tacit knowledge and workflows from people over time, rather than attempting to automate everything immediately.

How to start with autoresearch

Latent Space: What would you recommend to engineers who want to experiment with autoresearch?

Gavrilescu: The first step is to invest in your signals. What are the things you actually want agents to respond to?

Product feedback is a good example. Not all feedback carries the same value, and you cannot respond to every individual data point. You need a mechanism for filtering the signals and identifying which ones an agent should act on.

The second requirement is control over cost. You do not want to wake up to an unexpected thousand-dollar bill because an agent has been running an inefficient loop.

The third is to follow the research. Look at the kinds of harnesses models are being trained to use and remain close to those patterns. Study how research labs use data recipes and consider how those ideas can be applied to your own product.

The broader goal is to turn your product organization into a miniature research lab, with agents acting as miniature researchers.

  •  

How Cursor deploys AI inside the enterprise

Pauline Brunet, VP of Forward Deployed Engineering at Cursor, at AIEWF.

Forward deployed engineering has quickly become one of the most prominent roles in enterprise AI. Sitting somewhere between software engineering, product development and customer implementation, forward deployed engineers [FDEs] work directly with organizations to implement AI capabilities.

At Cursor, the role is especially ambitious. Pauline Brunet, the company’s VP of Forward Deployed Engineering, is building a team that works with organizations to implement agents across the entire software development lifecycle.

In an interview with Latent Space at the AI Engineer World’s Fair, Brunet discussed Cursor’s vision of an “AI software factory,” the challenge of expanding agent adoption beyond individual enthusiasts, and what engineers need to demonstrate if they want to move into forward-deployed work.

What forward deployed engineering means at Cursor

Latent Space: To begin with, how do you define forward deployed engineering?

Pauline Brunet: Forward deployed engineering depends on the business, the product, and the customer. You have to consider how configurable the application is. Is it something customers can use out of the box, or are you deploying something complex and highly configurable?

You also have to consider where customers are in their journey.

I don’t think of forward deployed engineering as a team that supports a traditional, out-of-the-box deployment. I think of it as a team that goes on-site, works inside a customer’s systems and tools, and deploys applications or platforms that help solve challenges at scale.

Those deployments are highly configurable and customized around the customer’s workflows, processes, systems, and tools.

Latent Space: Cursor’s customers are predominantly engineers. How does the FDE role apply to the way they use the product?

Brunet: Cursor is an AI coding platform and coding assistant. We work with people on AI-assisted coding, synchronous and asynchronous agents, and ultimately the idea of an AI software factory.

Today, we work with customers across many industries, including financial services, telecommunications, software development, technology, and semiconductors.

We help transformation leaders, IT leaders, and CTO organizations create an AI software factory across their operations. That includes how they plan and design software, how they write code, how they test and review it, and how they deploy and maintain applications at scale. So, very focused on the software development lifecycle from start to finish.

Building Cursor’s FDE team

Latent Space: How large is Cursor’s FDE team?

Brunet: We’re growing rapidly. Our goal is to grow the team tenfold by the end of December.

Latent Space: Are your current FDE employees primarily engineers, or does the team also include product specialists?

Brunet: They are all engineers. We hire software engineers with at least five years of experience and extensive customer-facing experience.

These are people who have developed and shipped code in production. They have built and designed systems, and they can make trade-off decisions and evaluate which systems or technologies should be used.

They also need customer-facing experience. We have people who previously worked at companies including Spotify, Rippling, and Palantir, and who have deployed production systems for customers.

From coding assistants to software factories

Latent Space: You mentioned the term “software factory,” which has begun appearing more frequently in the industry. What does that term mean to Cursor?

Brunet: For Cursor, it is about the software development lifecycle from start to finish: how you plan, design, write, review, test, and deploy code.

Today, those stages are often handled by different teams. You might have a design team, a development team, and a product manager working alongside them. Each group may be optimizing its own work with AI-assisted coding, but the process remains siloed.

We want to help customers across the entire lifecycle. You should be able to say, “Here is the feature I want to develop,” and then have long-running agents work with you across every step. That could include creating the plan and product requirements document, producing a demonstration of what the feature might look like, writing and testing the code, putting it into production, and maintaining it.

Issues and product feedback should also feed back into that same lifecycle. For us, a software factory means long-running agents helping people throughout that entire process.

Latent Space: So it is broader than agent orchestration alone?

Brunet: Correct. Exactly.

Moving beyond individual AI adopters

Latent Space: What problems are enterprises encountering as they try to implement agent technology?

Brunet: One challenge is that adoption is still concentrated among early adopters.

Within an organization, you might have 10% or 20% of people who are enthusiastic early adopters. They have done great work using local agents and cloud agents for their own tasks, and they have become highly productive.

What is missing in the next phase is the ability to use long-running agents across teams, processes, and workflows.

That requires more support from the top of the organization. Leadership has to say, “This is a priority, and this is how we want to automate or change this process.”

For the FDE team, it is therefore important to find the right champions inside an organization: people who want to meaningfully change the business and who will work with us and their internal teams to transform how work gets done.

Standardizing work with cloud agents

Latent Space: Local AI appears to be gaining momentum, partly because of the increasing availability of open-source models. Are you doing more local AI implementation work with customers?

Brunet: We have local agents that people run through the desktop application or the CLI, and that experience is largely self-service. People have adopted the technology at a phenomenal rate, particularly across Cursor’s user base.

We are also seeing people adopt cloud agents because they are excited about being able to run tasks without keeping their laptops half open. Agents can now work in the cloud on tasks that previously ran locally.

What becomes interesting is when this moves beyond an agent helping with one person’s job. The next question is how agents can work across a function, team, or organization so that processes are automated consistently. For example, you could have a QA agent applying the same process across several development teams.

We are receiving a lot of questions from customers about those kinds of use cases.

How customer deployments influence Cursor’s roadmap

Latent Space: Do the lessons from these deployments feed back into the core Cursor product?

Brunet: Yes. The forward deployed engineering team works very closely with customers on their use cases, so we are naturally a good way for the product and engineering teams to understand what customers want to build next.

We work closely with those teams and play a significant role in helping shape Cursor’s product roadmap.

The changing role of the forward deployed engineer

Latent Space: As agents become more autonomous, how do you expect the FDE role to evolve?

Brunet: I think the role is going to change drastically. I always say that if we are doing the same job we were doing six months ago, we have done something wrong.

Right now, people are still looking for inspiration about the use cases they can solve, so we want to propose new possibilities.

In software development, for example, we can show how designers and product managers might work seamlessly in Cursor alongside developers and testing teams.

We might also ask whether a company has considered using long-running agents to handle call-center or ticketing processes from start to finish.

As we work across industries such as healthcare, life sciences, the public sector, retail, and consumer packaged goods, we will continue identifying use cases across marketing, sales, and supply-chain operations. The FDE role will evolve alongside those possibilities.

How engineers can prepare for an FDE career

Latent Space: There are around 7,000 AI engineers at this conference. What advice would you give developers who want to move into forward deployed engineering?

Brunet: I’ve had this conversation five or six times already today. We are looking for builders with software engineering experience: people who have identified a problem and built a production-grade application or system from start to finish.

You should have designed it, developed it, tested it, and put it into production with real users.

My recommendation is to find those kinds of projects inside your organization and take ownership of them from beginning to end. Make sure you understand why you made each design decision.

How did you select the database? How did you choose the different services? Why did you design the system in that particular way? What were the trade-offs?

You should also understand the measurable return on investment, both in traditional business terms and through evaluations that demonstrate the value you are creating for internal customers.

If you want to get into forward deployed engineering, become familiar with these kinds of projects, gain experience delivering them, and learn how to explain the decisions you made.

  •  

Warp CEO Zach Lloyd on why software factories are the next phase of coding

Warp founder Zach Lloyd in the AI Engineer World’s Fair expo hall.

I’ve been covering Warp for a couple of years now, and its rapid evolution from a command-line interface tool to a software factory platform has been fascinating to watch. The company began in the pre-ChatGPT days, in mid-2021, as a Rust-based terminal. Then when AI hit, it turned into a terminal with integrated coding agents.

But the competition among CLI tools has dramatically increased in recent years, including from Claude Code, Codex CLI, and Gemini CLI — three products backed by massive tech companies. This likely led to Warp’s decision to open-source its core CLI tool in April this year.

I’m a Warp user myself, finding it a much more sophisticated tool than my native Mac CLI. But I also admire the company’s ability to adapt to the times — a trait I spotted in CEO Zach Lloyd during my first interview with him a couple of years ago. So I was keen to catch up with him at the AI Engineer World’s Fair this week, where he presented a keynote session on software factories, the new term for orchestrating a team (ahem, a factory) of agents.

Warp has a new agent orchestration platform called Oz. It’s the company’s answer to what Lloyd believes is an industry transition, from engineers working interactively with agents to automated systems that continuously triage, implement, review, verify and monitor software changes. Oz is intended to connect multiple models and coding harnesses across local environments and isolated cloud sandboxes, while fitting into tools developers already use.

I spoke to Lloyd just after he made his presentation on-stage, which you can view on YouTube — it’s a good primer to what software factories are. In our one-on-one discussion, we get into the reasons Warp made its software factory pivot, how Lloyd came up with the term (independently, it seems, from similar companies — like Factory), and why he expects most significant software projects to operate some form of automated factory within the next year.

From individual agents to an automated development loop

Latent Space: When did you first come across the term “software factory,” and what attracted you to the concept?

Zach Lloyd: I can’t remember exactly when I started conceiving of it in those terms, but it was within the last six months, as the ability to automate software development became more complete.

We started with more one-off automation: run an agent in the cloud. A lot of platforms began there. Then it became: run an agent in the cloud on a timer.

The next question was, what is the most valuable loop to automate? The answer is basically the main loop of software engineering: triage, specification, implementation, review, verification, shipping and monitoring.

We began building toward this cloud-automation vision about a year ago, before we started building Oz. Over the past few months, the industry has also begun coalescing around the ‘factory’ term. There is an entire software-factory track at this conference.

It is literally what we are gearing our product around. In the next version of Oz, you will set up your factory, see what it looks like and manage the factory floor.

But I don’t care that much whether the term sticks. The essential shift is from interactive development to automated development. “Factory” is a useful metaphor for that.

Building the factory around existing workflows

Latent Space: In your presentation, you showed a software-factory stack containing several of your own products. Is Warp’s plan to provide the tools that make up that stack?

Lloyd: Yes. When you enter Oz, our cloud-agent platform, you will be walked through setting up a factory.

You choose your repositories, the parts of the software lifecycle you want to automate, and the points where humans should be brought into the loop. Different organizations and codebases will have different preferences. Do you fully automate code review? Do you have humans review certain high-risk changes?

The system then starts creating the loop. It might pull issues from Jira or Linear, let people submit them through Slack or Teams, and allow developers to redirect an agent from GitHub.

What is interesting from a product perspective is that most of the factory is not necessarily a new interface. It is an integration into people’s existing workflows. That is how we are conceiving it, at least.

Why Warp is moving beyond the terminal

Latent Space: When I first wrote about Warp, it was building a modern terminal. Code is still important now, but increasingly it is being produced by agents. It looks like Warp has broadened its product vision accordingly...

Lloyd: One hundred percent. A good way to think about it is that the company’s mission has stayed the same since we founded it. It has always been about empowering developers and companies to ship better software more quickly.

The product has evolved tremendously. It began as a modern version of the terminal, before the current AI wave. The next iteration was a terminal with agents built into it, which we are still investing in and which we have now open-sourced.

But the world keeps changing. The underlying AI improves so quickly that my view of the future is what I described in the talk: the interactive component is going to become less important.

As a company, you will want a central place where software gets built and where you can measure the efficiency of that process. I’m not afraid to redirect what the product becomes. As the underlying technology gets better, companies that do not adapt are going to be left behind.

Factory engineering as a new discipline

Latent Space: The word “factory” may be off-putting to some developers, given its connotations with mechanism and rote work. What feedback have you received from AI engineers about this pivot?

Lloyd: The concept resonates strongly with the economic buyer — the person running the engineering team.

For an individual engineer, it can sound mechanized and uncreative. They may think: “I enjoy coding. Why would I want to work in a factory?”

One point I tried to communicate in the talk is that this will become a new engineering discipline. I think it can be extremely interesting if you view the job as meta-engineering: building the system that builds the product.

It uses many of the same problem-solving skills. You are asking why an agent performs one task well and another poorly. How should you adjust its feedback? What context does it need? How should the workflow change?

But, for better or worse, the power of these systems and their ability to accelerate software development are so great that writing everything by hand is not going to make sense for much longer.

Where forward-deployed engineers fit

Latent Space: Another trend at the conference is forward-deployed engineering, which often combines aspects of product management, consulting and traditional engineering. How does that fit into the software-factory model?

Lloyd: Standing up a software factory potentially involves integrating with a large number of existing systems, depending on the company.

The factory will work most effectively when it has context from those systems and is integrated throughout the organization’s workflow. A lot of forward-deployed engineering work in this area is effectively a transformation project.

It requires real engineering from someone who understands how to configure and deploy one of these systems. We do some of that, and some of our competitors do as well.

I don’t know what the final state will look like. Warp is approaching it more as a platform business than a services business. But there is certainly a business today in sending smart people into a company to transform its workflow using these products.

Warp as the test bed for Oz

Latent Space: I use Warp as my terminal, including for some coding tasks. What happens to the original Warp CLI product in the software-factory era?

Lloyd: When we open-sourced Warp, we put the repository under the control of Oz. We built a software factory around the open-source project, using our own factory platform.

We are still trying to improve Warp as much as possible. We are doing it with the community, and we are doing a lot of it with agents. In that sense, Warp is a test bed for the factory concept.

But it is also a product used by almost a million developers, many of whom rely on it as their primary development environment. We use it constantly ourselves, and we still have internal engineers whose job is to improve it. We are simply approaching that work with a factory mindset.

Gradual automation, not an overnight replacement

Latent Space: What do you expect the next year to look like, in terms of adoption of software factories?

Lloyd: This will not happen all at once. Engineers are not going to wake up one morning and discover that a software factory has replaced their jobs.

Companies will start with specific use cases, certain types of issues or lower-risk repositories. Those are places where they may be comfortable not having a human review every single line of code.

They will see how it performs. Then the engineering challenge becomes: instead of merging 20% of pull requests automatically, can we get to 30%, 40%, 50% or 60%?

There will still be a remaining percentage of work done by people because it is too difficult, ambiguous or dependent on greenfield thinking.

But I think this shift will happen over the next year. My prediction is that every significant software project will have some engine of code — something resembling a factory — continuously driving it forward.

It will become similar to GitHub or CI/CD: a standard part of how serious software projects operate. I would be surprised if that did not happen.

Start by automating the annoying parts

Latent Space: There are thousands of AI engineers at this conference. What should they do to prepare for this shift?

Lloyd: Instead of only building the product directly, try building some automation toward a factory and see what it feels like.

Suppose you want an agent to implement incoming user issues automatically. What is involved in making that work? What prevents you from adopting it?

Perhaps code review is the bottleneck. Perhaps the agent is making changes, but you cannot clearly see what it did. You only discover those problems by trying to build the loop.

Get out of the mindset of building everything by hand. Find an annoying part of your job and try to create a loop that handles it for you using a factory approach.

  •  

AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers

Agents are here to serve you in the software factory.

Loops, loops and more loops. That word, loop, dominated conversations on day 2 of the AI Engineer World’s Fair — the first full day of keynotes and sessions. Perhaps knowing in advance what everyone would be talking about, AIEWF cofounder swyx titled his opening talk, “Loopcraft: The Art of Stacking Loops.”

swyx began by commenting on the evolution of AI engineering from 2022: from chat, to tools, to goals. “These days, we’re all about automations,” he added. “We’re all about cron jobs and loops.”

Allie Howe, a member of technical staff for Keycard, then introduced the main stage track for the day: Software Factories. She referenced Geoffrey Huntley’s influential article, “everything is a ralph loop,” a theory about turning an AI coding agent into a persistent worker by repeatedly restarting it against the same spec.

Pablo Castro from Microsoft then talked about Foundry, the company’s “AI app and agent factory.” He claimed that a “learning loop” occurs when people and agents work together.

OpenAI’s Alexander Embiricos and Romain Huet were next on, and they focused a lot on Codex, the company’s coding agent. One point they made was that using multiple agents via loops can result in enhanced productivity.

“There will be a lot of talk today about loops,” Embiricos said. “And if you can connect the agent to not only the work that you have to do, but why it has to be done, that’s how you can get the agent to start to begin much more work. And then if you can connect it to what you do afterwards, review and deploy, that’s how you help it land much more work.”

This segued to a presentation by Peter Steinberger, the “ClawFather” of OpenClaw, now working for OpenAI. He too was all-in on loops, noting that he designs loops to manage agents. He added that deciding what to pay attention to is his main challenge nowadays — and that the future is “better loops” to help solve this issue.

Software factories

All this talk of looping led naturally to the concept of “software factories,” the subject of a presentation by Tereza Tížková from a company called Factory. She defined a software factory as “the whole loop, the whole lifecycle of developing software with autonomy.” She added that this doesn’t mean just coding, but also “collecting all the signals, reacting to user feedback [and] to logs, prioritizing what’s important, then orchestrating it all.”

Zach Lloyd from Warp also spoke about software factories; in fact, his thesis was that “software engineering will become factory engineering.” Loops in Lloyd’s framing were about improving the system.

In both Tížková and Lloyd’s talks, the emphasis was on having the agents doing the building for you. “You’ll be building the thing that builds the product,” was how Lloyd put it.

Afterwards, I went down to Warp’s booth in the AIEWF expo hall and spoke to Lloyd about software factories. I particularly wanted to know why Warp, which began as a CLI tool for developers, has pivoted into a ‘software factory’ platform where developers aren’t supposed to do coding anymore.

“The way to think of the factory is, like, pick your repos, pick the parts of the lifecycle that you want to automate, pick the ways in which you want humans to be brought into the loop,” Lloyd told me. “And different organizations [and] code bases will have different preferences for, like, do you fully automate code review [or] do you have humans do hard coding, stuff like that.”

I noted that the term “factory” might be offputting to many developers, since it implies mechanized rote work — much different from the creative era of coding we’ve just come from. Lloyd recognizes this is a challenge, but he argues software factories will become a new discipline of engineering — and that it still requires problem solving.

“For better or worse, the power of these systems is so great and the ability to accelerate is so strong that just writing stuff by hand...I don’t think it’s going to make sense for very much longer,” he said.

(For more from Zach Lloyd on software factories, stay tuned for a Latent Space interview to publish shortly.)

Forward Deployed Engineers

Related to loops and software factories, another theme from AIEWF today was the trendy new role of Forward Deployed Engineers. In an interview with Natalie Meurer, Head of Agent Engineering at Sierra, I established that FDEs are also sometimes called “agent engineers.” The main point is to help organizations adapt to agents, from a development perspective.

Meurer pointed out that a lot of the work of integrating AI into companies these days is in orchestrating agents.

“In practice, most customer-specific work takes place at the orchestration layer rather than in the models themselves,” she told me.

Cursor’s VP of Forward Deployed Engineering, Pauline Brunet, also ran a session today at AIEWF, in which she positioned FDE as part of the shift to software factories. “We partner with your organization to co-design and co-build your AI software factory,” she said. “We transform how you design, develop, and maintain software across your entire life cycle.”

(More insights from Brunet coming in an upcoming Q&A.)

Open Source AI

Another key theme from AIEWF today was the rise of open source AI. Zixuan Li, the head of intriguing new Chinese company Z.ai, was due to make an appearance at the conference. Because of travel issues, he couldn’t make it in person. He did make a virtual presentation, though, focusing on the company’s groundbreaking open LLM, GLM-5.2 — its “flagship model for long-horizon tasks.”

He also introduced ZCode, a harness that “supports all frontier models.” Li compared it specifically to OpenAI’s Codex.

HuggingFace’s Thomas Wolf then interviewed Olive Song from Chinese company MiniMax, which recently released its latest open-weight model, M3.

Open source AI is a big reason why local AI is becoming more popular. Ahmad Osman is the founder of Osmantic, a company building open source software for deploying and operating local AI systems. He spoke to us today and noted that open models have improved dramatically in recent times.

“Architectures are becoming more efficient, and many small improvements compound,” he said. “Once a frontier lab demonstrates that a capability is possible, the open source ecosystem can work backwards from that and find ways to reproduce it more efficiently.”

Conclusion

Those were the big trends from day 2 of the AI Engineer World’s Fair. I’ll be back tomorrow with all the action and analysis from day 3. Don’t forget to tune into the keynotes on YouTube if you’re following from work or home.

  •  

Forward Deployed Engineers and the future of software engineering

Sierra’s Natalie Meurer at the AI Engineer World’s Fair today.

Natalie Meurer is Head of Agent Engineering at Sierra, where she leads a global team of more than 120 engineers building conversational AI agents for enterprise customer service. Before joining Sierra, she worked in technology policy, taught herself to code and spent five years at Palantir.

Forward deployed engineering (FDE) was one of the tracks running at today’s AI Engineer World’s Fair. As Meurer explained to Latent Space before the session she presented, FDE began as a model for placing highly technical employees close to customers. But the title now covers a wide range of roles across the AI industry — including what Sierra calls the agent engineer: an engineer who combines systems integration and agent development with an understanding of customer operations, product, and the end-user experience.

In this Q&A, Meurer argues that FDE is defined more by accountability than by a particular skill set, adding that product and customer-facing engineering may be starting to converge.

Defining forward deployed engineering

Latent Space: What is your definition of a forward deployed engineer?

Natalie Meurer: That is really the point of my session: the role lacks a consistent definition.

If you look at its historical trajectory through to the present, it is more clearly defined by accountability to customers than by the shape of the role or the work you are doing.

There is power in having that accountability. But the range of associated skill sets has become so broad that it can almost become nonsensical.

Latent Space: How did you get into this kind of role?

Meurer: I began in technology policy. I was a policy nerd who learned to code on the side, which earned me a role as an engineer on the privacy team at Palantir.

I spent about five years there, working across law enforcement, defence and infrastructure engineering. I then went to business school because I wanted to bring the business dimension into the mix. After that, I joined Sierra and founded the agent engineering function.

Why Sierra calls them agent engineers

Latent Space: Did Palantir’s forward deployed engineering model influence the role at Sierra?

Meurer: Somewhat, although we intentionally called the role agent engineer, rather than forward deployed engineer.

Forward deployed engineering can mean so many things. We thought the title should capture the shape of the technical work, rather than only the customer-obsession element. That is why we chose agent engineer.

I see agent engineering as either a subset of, or adjacent to, forward deployed engineering. It describes a more specific form of customer-facing engineering focused on developing agents.

What an agent engineer does

Latent Space: What does your team do when working with a customer?

Meurer: Sierra builds conversational AI agents for inbound and outbound customer service. Our work includes integrating customer systems with low-latency voice and chat agents, as well as agents that operate over email.

The role requires technical skills such as data integration, but it also requires taste. You need to understand what sounds good and what will feel human when you are designing a voice agent. That element is particular to agent engineering.

Latent Space: Does an engagement begin with a defined use case, or do you help the customer decide what to build?

Meurer: We conduct discovery with our customers. We try to find the intersection between problems that are genuinely difficult — because we are good at difficult problems — and problems that will have a meaningful business impact.

In financial services, for example, that might begin with dispute processing. It is complex and needs to be done correctly, but it is also a high-emotional-intelligence interaction. If somebody sees a fraudulent charge on their credit card statement, they may be frightened, and the agent needs to calm them down.

Almost every Sierra customer is also somewhere on the trajectory towards using an agent as its front-door interactive voice response system: the first entity that answers when a customer calls.

The hard work is often above the model layer

Latent Space: How much of the work involves the underlying AI models?

Meurer: We think of our agents as an orchestrated constellation of models. Internally, we are constantly evaluating the best model for a particular job, and we bring the best of that work to our customers.

In practice, most customer-specific work takes place at the orchestration layer rather than in the models themselves. We sometimes integrate with a customer’s own models, and we also help customers use the platform and build agents themselves.

A lot of the work involves helping them apply their internal knowledge and context.

Custom deployments and reusable patterns

Latent Space: How much of the work is customer-specific, and how much can be reused?

Meurer: It is a mixture of both.

Every customer is building an agent that is intentionally specific to its organization. It should represent the best possible interaction with that particular brand.

Other capabilities are more reproducible. Answering questions from a knowledge base, for example, is a fairly universal problem. We also have industry experts across financial services, healthcare, travel and hospitality, and retail who bring domain knowledge and best practices.

But the fundamental appeal of what we are selling is something custom. We have seen large organizations across industries reach production in as little as 40 to 60 days.

Each agent is still customized around the customer’s APIs, systems, standard operating procedures, brand and tone.

Agents as enterprise systems

Latent Space: Is agent development becoming primarily an orchestration problem?

Meurer: There are many different flavours of multi-agent architecture. The term “agent” can refer to the entity that answers the phone, but it can also refer to a sub-agent or even a single prompt equipped with tools.

Every enterprise we work with wants to know how it can maintain everything its agentic ecosystem is capable of doing. It needs to manage all the integrations and all the teams that contribute to the agent.

Part of that is a change-management problem.

At Sierra, we tend to think of a single agent as managing the entire customer interaction, regardless of the particular subtask involved. We call those subtasks journeys.

Enterprises nevertheless need a way for hundreds or thousands of people to contribute to these systems, understand what is changing and follow a discrete release process.

Product engineering and FDE are converging

Latent Space: As companies develop more internal expertise, how will the FDE role evolve?

Meurer: I think it will remain customer-facing. But when code becomes cheap to author, it also becomes easier to translate customer insights directly into a product.

Product engineering and forward deployed engineering are therefore converging in some respects — at least among the best people in each role.

If you are a product engineer, you should be talking to customers. If you are a forward deployed engineer, you should be building the product. I think that is new.

Being customer-facing will remain important. Even if you had an AGI-like reasoning model that could work out how to perform a process each time, you would still need to encode that process appropriately.

You do not want the system independently figuring out how to handle an order return for the 100,000th time that week. You want a consistent process that it follows.

That makes customer service different from some other agentic use cases. A coding agent is often trying to solve a new problem for the first time. In customer service, you are solving essentially the same problem, framed slightly differently, perhaps 100,000 times a week.

That creates a different need for both the platform and the partner helping the customer encode its rules. Agents will become easier to build, but there will always be a place for people who can work with customers and translate what they learn into the product.

Why generalists may become more valuable

Latent Space: Will developers increasingly need product and customer-facing skills?

Meurer: That is my belief. I think the best developers will develop those skills.

Many people are asking what the engineering role will look like in one or two years. One view is that specialists will become even more important because they possess knowledge that is not readily available to an agent.

The other view, which I lean towards, is that generalists will become more valuable.

Forward deployed engineering has historically been the classic generalist role because it combines engineering with the customer-facing nature of the job.

Forward deployed engineers — or agent engineers — therefore inhabit one of the most forward-looking areas in AI and engineering.

Latent Space: Could “agent engineer” eventually become the default term?

Meurer: I am not sure. I expect engineering as a whole to move towards a more holistic definition, one that may incorporate more of what we currently call forward deployed engineering.

The market currently has go-to-market engineers, forward deployed engineers, agent engineers and AI engineers.

I think all of those will become different parts of the engineering craft. We will also discover entirely new jobs for engineers to do.

  •  

Ahmad Osman on why local AI is catching up

Ahmad Osman at AIEWF
Ahmad Osman at the AI Engineer World’s Fair today.

Ahmad Osman has been advocating for local AI — running models on your own computer, workstation or dedicated hardware — long before it became a major theme at this year’s AI Engineer World’s Fair. He is also the founder of Osmantic, a company building open source software for deploying and operating local AI systems.

One of the themes emerging from AIEWF is that open source LLMs are becoming increasingly credible alternatives to large, proprietary frontier models. Since most local AI systems depend on open models, that shift strengthens the case Osman has been making. As he told Latent Space, “the gap between open-source models and closed-frontier models keeps shrinking.”

Osman makes the argument even more explicitly on a website called Open Source AI Must Win, where he writes that “the ability to study, build, repair, deploy, audit, adapt, teach, preserve, and run intelligence systems without asking permission is of existential importance.”

At AIEWF, Osman ran a two-part workshop on local LLMs and workstation agents. The sessions showed how quickly the field is moving — from models running on phones and laptops, to dedicated GPU workstations and enterprise infrastructure.

The interest in Osman’s workshops was not limited to hardware hobbyists, either. Attendees ranged from students considering their first AI-capable machine to enterprise executives thinking about model routing, private infrastructure and control over company data.

In the following Q&A, Osman explains why local AI is attracting more attention, how the model and hardware landscape has changed, and why he expects more developers and enterprises to begin treating local AI as serious infrastructure.

Making local AI tangible

Latent Space: Can you summarize what the workshops were about and what attendees were looking for?

Ahmad Osman: It was a two-part workshop, and there was more demand than we had space for. Some people unfortunately had to be turned away.

I came in with a website we had prepared to demonstrate local AI. It was essentially a hardware arena where people could compare systems such as the DGX Spark, AMD Strix Halo machines and other devices. You could run them against one another, or compare them with a frontier cloud model, and see the performance, output quality, speed and latency for yourself.

The main idea was to make local AI feel real. There is still a perception of it that dates back to 2022, when the models were much less capable. But everything has improved substantially since then.

There is still a lag behind frontier models — perhaps four to eight months — but local and open models are catching up. We wanted people to interact with these systems rather than just hear a theoretical argument about them.

The software behind the demo is open source and available on GitHub. The second workshop went further into setting it up and showing the full system in action.

A model is only one part of the system

Latent Space: What is missing when people think of local AI as simply running a model on their own machine?

Osman: There is a big misconception about products such as ChatGPT or Claude Code. They come with a complete infrastructure around the model and around the agent. It is not just one thing.

A friend of mine bought an RTX 5090 to run Qwen 3.5 locally. He connected Claude Code to the model and asked it to change the RGB lighting on the GPU, but it failed. He then used the hosted Claude Code service, and it worked.

I asked whether he had given the local model internet search access. He had not. The model’s training data had a cutoff date, while the software and documentation he needed had since changed.

Once we gave the local system access to a search endpoint, it was able to complete the task.

That is the point: when you use a hosted agent, you are not only using a model. You are using search, tools, infrastructure and other services around it.

With our open source deployment system, we are trying to provide the complete experience — from a chat interface and document ingestion to agents, harnesses and search tools. That end-to-end layer has been lacking in the local AI ecosystem.

Interest spans students, enthusiasts and enterprises

Latent Space: Who came to the workshop? Were they mainly hardware enthusiasts, or people trying to build privacy-based applications?

Osman: It was a very wide audience.

At the end of the second workshop, a student asked me what hardware she should buy before going to college. An executive from Intel asked how we could get the software running on Windows in a particular way to improve the user experience.

Some people were enthusiasts. Others had very enterprise-focused questions. The common thread was interest in running something they can control, whether that means a model on a MacBook, a GPU at home or a dedicated cluster of high-end enterprise hardware.

People asked about enterprise model routing, data collection, traces, agent sandboxing and latency. Others asked how many GPUs I have at home. The answer is 22 RTX 3090s.

The breadth of interest surprised me. This was my first AI workshop, and I was lucky enough to do two of them back to back.

You may not need to buy a GPU

Latent Space: Do developers need to go out and buy GPUs to experiment with local AI?

Osman: It depends on the size of the model you want to use.

You can run a four-bit Qwen model on a MacBook. At the other extreme, a very large frontier-class open model might require several RTX Pro 6000 GPUs.

But the broader trend is that models are becoming much more efficient. On a modern phone, you can now run a model that outperforms systems people were using in the cloud only a couple of years ago, without using all of the device’s memory.

That shows how far model efficiency has come in a relatively short time.

Models and hardware are improving together

Latent Space: Is the progress mainly coming from better software and models, or from hardware as well?

Osman: The models have improved dramatically.

Architectures are becoming more efficient, and many small improvements compound. Once a frontier lab demonstrates that a capability is possible, the open source ecosystem can work backwards from that and find ways to reproduce it more efficiently.

We are seeing models with tens of billions of parameters deliver performance that would previously have required much larger systems. Some of those models can run on an RTX 3090 released in 2020. Two years ago, that level of capability on that hardware would not have been realistic.

This is still a very new field, and we do not know the end state. But we know the systems will continue to improve.

The rise of hybrid and sovereign AI

Latent Space: Do you expect more applications to combine local and cloud AI?

Osman: Yes. Edge models are going to become more popular, and this is not only about consumers.

Enterprises are increasingly aware that the models they depend on may not always remain available to them in the same form. Providers can change quality, pricing, access or policies.

That creates an incentive to move toward dedicated hardware and secure compute. It does not necessarily have to sit on premises. A company can use dedicated, colocated hardware that it controls.

The benefit is that the quality of the model does not unexpectedly change, access cannot simply be removed, and the company retains control over its intellectual property, data, privacy and compliance obligations.

Open source models are also continuing to close the gap with frontier proprietary systems. We have seen a rapid progression through Llama, Mistral, Qwen, DeepSeek, GLM and Kimi models. Each generation narrows the gap.

Specialized models may be the real opportunity

Latent Space: Where do you think this leads for businesses?

Osman: I have believed for some time that smaller, specialized models are the future for many business use cases.

An enterprise may begin with a general model and collect traces, messages and feedback from how employees use it. Over time, that data can support a more specialized model tuned to the company’s particular work.

That can improve performance, reduce costs and make the system more useful for the business.

I also think open source model companies may increasingly monetize through licensing for fine-tuning, reinforcement learning or specialized commercial deployments.

As more companies move away from relying entirely on cloud APIs and secure their own compute, these labs will have an incentive to keep releasing strong open models while capturing value when businesses adapt them for proprietary use cases.

The broader direction is toward greater sovereignty: companies and individuals controlling their models, compute and data, while still benefiting from the rapid progress of the open source ecosystem.

  •  
❌