❌

Normal view

SpaceXAI's Grok Bot turns agents into persistent digital coworkers that can operate your apps for $120-per-month

SpaceXAI, the division of SpaceX formerly known as xAI, is launching an early beta version of Grok Bot, a new agent designed to move AI assistants beyond answering prompts and toward continuously executing work across the software employees already use.

The central idea is straightforward: instead of opening an AI assistant whenever a task arises, users create persistent Bots with specific jobs, give them access to applications and websites, and delegate work much as they would to a teammate.

Each Bot operates through its own computer environment, can continue working when the user's laptop is closed, and can return when it needs approval or has finished the assignment.

SpaceXAI says the system began as an internal prototype before spreading across the company, where teams created Bots for sales outbound, marketing campaigns, office operations, bug fixes and other work. The company is now turning that internally developed workflow into a product for external users.

β€œBots are AI teammates that do real work for you,” the company said in announcing the product. β€œThey sign in to your tools, use them just like you do, and come back with finished work.”

The company did not release benchmarks for Grok Bot's performance on agentic tasks. And it arrives amid an increasingly crowded marketplace of first-party AI agents that attempt to reliably complete real, enterprise workflows by interfacing with a user's other applications and devices.

Anthropic introduced computer use for Claude in 2024, allowing models to inspect screens and operate interfaces through mouse and keyboard actions, and continued expanding with the launch of the developer focused Claude Code harness in early 2025 and the more non-technical, white collar focused Claude Cowork agent early this year.

Meanwhile, OpenAI gave its Codex harness the ability to control other computer apps in April, launched agentic Workspace Agents that can also connect to third-party applications and use them autonomously, and recently debuted a new ChatGPT Work environment for longer, multi-step tasks and finished deliverables.

Grok Bot seeks to join the party with its own management model for agents: persistent workers with responsibilities, memory, learned routines and the ability to hand work to one another.

Pricing and availability: Grok Bot starts at $120 per seat per month for teams, $200 per month for individuals

Grok Bot is available beginning today, August 11 in beta for SuperGrok Heavy, Cursor Ultra and Cursor Premium Teams subscribers (recall SpaceX acquired Cursor for $60 billion back in June). The product arrives for macOS, Windows, Linux and iOS, with Android listed as coming soon.

According to its product page on xAI.com, Grok Bot is included with Cursor Ultra at $200 per month for individuals. The plan includes a computer for Grok Bot, access to users' tools, scheduled routines, desktop and mobile operation, and extended AI-token limits.

For organizations, Cursor Premium Teams costs $120 per seat per month and adds centralized billing and settings, a team marketplace for skills and plugins, shared usage analytics and SAML/OIDC single sign-on.

Existing SuperGrok Heavy ($300 per month) subscribers also receive access. However, for organizations wishing to sign up today, SpaceXAI is directing them to a waitlist for future access.

Those prices make Grok Bot a substantially different purchasing decision from a low-cost general AI subscription. The economic question for companies will be whether persistent Bots can replace enough manual work or conventional automation infrastructure to justify the per-user cost β€” and how usage limits affect total cost once agents begin running continuously.

From prompting an AI to managing one

SpaceXAI describes Grok Bot as a team of β€œalways-on agents.” Users can create multiple Bots, assign each a role and let them work simultaneously.

The company provides examples including Sales Outbound, Talent Scout, Paid Media, Expense Manager, Product Performance, Bug Reproduction, Account Health and Chief of Staff. A sales Bot, for example, can research accounts, score prospective contacts, prepare email and LinkedIn outreach in the user's voice, and assemble the results for human approval.

Promotional materials show SpaceXAI using the system internally for substantially longer chains of work. One sales Bot can add call-transcript notes to a CRM and draft follow-up messages. An operations Bot can seat new hires and process invoices arriving through Gmail. An engineering Bot can reproduce a bug in the product interface, file a ticket and then hand the repair to a debugging Bot.

The architecture could make Grok Bot particularly relevant for workflows that span systems that were never designed for AI automation.

Rather than requiring every application to expose an API specifically for an agent, Grok Bot can sign into applications and websites and operate their interfaces. SpaceXAI says Bots have their own computers and can continue working 24/7.

The company explicitly says this includes websites and applications that have β€œno clean API or MCP,” an important distinction for enterprises with legacy software, fragmented SaaS environments or internal systems that have never been instrumented for agent access. Instead of limiting automation to formally integrated services, Grok Bot is designed to work through the same software interfaces a human employee would use.

The company says early users are already applying Bots to jobs including vendor negotiations, e-commerce customer support and continuously updating CRM systems.

Another feature attempts to reduce the engineering required to automate repeatable business processes. Users can demonstrate a workflow while a Bot follows along. Grok Bot can then save the process as a routine and execute it later without requiring the user to reproduce every instruction.

SpaceXAI says the Bot can also incorporate corrections into those learned routines, allowing the workflow to change as the user teaches it how a particular process should be handled.

That potentially changes the deployment model from explicitly programming an automation to teaching an agent how an employee performs the job.

The company is also claiming a more persistent form of behavioral memory than simply retaining a chat transcript. According to the launch announcement, Bots remember prior conversations, learn preferences such as a user's writing voice and edge cases, and gradually learn when they should interrupt for approval versus continue independently. SpaceXAI says they can later resume dropped threads, nudge stalled handoffs and pick up work from earlier conversations.

It further says Bots can become proactive over time, sometimes identifying work before the user explicitly asks for it. That is a more ambitious claim than conventional scheduled automation and will put additional pressure on permission controls and escalation rules if the system is deployed against production applications.

Bots can delegate work to other Bots

Grok Bot also supports multiple agents operating together.

Users can place several Bots into the same thread, where the agents can pass work between one another. The company's demonstration includes specialized Research, Communications, Chief of Staff and Travel Bots coordinating tasks.

SpaceXAI says those Bots can independently message one another and share context within threads. Users can also put multiple Bots into a group conversation where they assign ownership, transfer work and coordinate among themselves, bringing the human back in primarily for judgment calls.

Internally, the company says employees sometimes place a Chief of Staff Bot above specialist Bots responsible for functions such as inbox management, recruiting, expenses, operations and bug fixes. That makes the product's orchestration model more explicit: the user does not necessarily have to serve as the routing layer between every specialized agent.

Initial reactions are extremely positive

Lenny Rachitsky, host of the popular vlog and podcast Lenny's Podcast and author of newsletter Lenny Letter, received early access to Grok Bot and loved using it. As Rachitsy wrote on X : "I haven't been this excited about a new AI product in a while. It's like OpenClaw, but super easy, reliable, and less scary to use. I think this will be a huge new product line for Cursor/Grok/SpaceX."

Similarly Matt Shumer, an AI entrepreneur who said he tested Grok Bot for several weeks before launch, highlighted this orchestration as one of the product's strongest features.

β€œThe best way I can describe it is an agent for everything, not just code,” Shumer wrote on X.

In one test, Shumer said he created separate researcher and writer Bots, then created a Chief of Staff Bot and instructed it to coordinate the other two on a project. He expected the workflow to break down.

β€œIt worked out of the box,” he wrote.

His main criticism involved model selection.

Unlike systems where developers or advanced users explicitly select the underlying model, Shumer said Grok Bot automatically routes tasks to models on the backend.

β€œYou don’t choose a model for your Grok Bot,” he wrote. β€œIt’s all done automatically on the backend.”

Shumer said the model router β€œwasn’t great” during his testing, although he said he was subsequently told it had improved.

SpaceXAI's expanded announcement still does not identify which underlying models the router uses, nor does it document a mechanism for users to select, pin or switch to a particular xAI or third-party model. As a result, the model layer remains largely abstracted from users in the publicly supplied launch material.

That abstraction represents an important tradeoff for enterprise deployments. Automatic routing can remove a significant configuration decision for ordinary employees, but advanced users may want explicit control over model cost, latency, reliability and behavior β€” particularly for repeatable production workflows.

The agent market is moving toward longer-running work

Grok Bot enters a market increasingly focused on agents that can do more than generate text or code.

Anthropic's computer-use capability established a mechanism for Claude models to interact with software through screenshots, cursor movements, clicks and typing. Its broader Claude product also connects with workplace services and remote MCP servers.

OpenAI, meanwhile, now describes ChatGPT Work as an agent for β€œlonger, multi-step work and finished deliverables,” while keeping Codex focused specifically on software development. OpenAI's enterprise agent economics can also incorporate usage-based credits, making task complexity and token consumption part of deployment cost calculations.

Grok Bot's differentiation is therefore less about proving that AI can operate software than packaging computer use, persistence, workflow learning and multi-agent coordination into something resembling a workforce interface.

SpaceXAI's announcement sharpens that distinction by emphasizing completion rather than assistance. One company product employee, identified only as Roman, describes the difference as closing the gap between work that is nearly finished and work actually completed inside the destination application: β€œGrok Bot can finish the swing, because the work lands where a human would put it, in the actual tool.”

That distinction will ultimately depend on reliability. A chatbot producing a bad answer creates a correction problem. An autonomous agent operating CRM records, support queues, vendor conversations or other production systems can create an operational problem.

Grok Bot's success will therefore depend not only on model intelligence, but also on permissions, predictable execution, escalation behavior, memory accuracy and how reliably agents recognize when human approval is necessary.

That challenge becomes more significant if Bots act proactively, resume forgotten work and coordinate with one another without the user serving as an intermediary. Those capabilities reduce the amount of supervision required when they work correctly, but they also expand the consequences of an incorrect assumption, stale context or improperly scoped permission.

The interface may matter as much as the models

Shumer described the product's interface as feeling like iMessage, an intentionally familiar metaphor for a system whose underlying architecture β€” autonomous computers, persistent memory, agent orchestration and automatic model routing β€” could otherwise be difficult for nontechnical users to configure.

SpaceXAI makes essentially the same usability argument in its launch announcement. Rather than asking users to construct workflows before getting started, it says users can simply message a Bot from a phone or desktop, hand it work and later continue the same conversation from either device.

That simplicity is part of the product strategy. Grok Bot is trying to hide much of the conventional machinery of automation β€” workflow builders, explicit integrations, agent routing and orchestration β€” behind an interaction model that resembles messaging a coworker.

That may prove to be the larger bet behind Grok Bot.

The AI industry has spent several years making models increasingly capable of using tools and completing multi-step tasks. Grok Bot attempts to turn those capabilities into an organizational abstraction people already understand: give someone a job, teach them how you work, and let them coordinate with the rest of the team.

If that abstraction proves reliable, the enterprise agent competition may increasingly shift away from which assistant produces the best individual response and toward which platform can most reliably manage fleets of agents performing ongoing work.

Why AI-driven purchase intent so rarely becomes a completed sale

11 August 2026 at 15:00

Presented by Rezolve Ai


When an AI assistant recommends a product or brand, it generates something valuable: a purchase-ready consumer with high intent and low friction in their decision. That consumer has already compared options, asked follow-up questions, and arrived at a conclusion. They want to buy.

What they encounter next is a commerce infrastructure that was not designed for them.

The gap between recommendation and purchase

The typical enterprise commerce stack was built for a specific model: a consumer who arrives at a brand's website through search or a direct link, navigates product pages, adds to cart, and completes checkout through a multi-step form flow. That model assumed the consumer would do the work of bridging their intent to the transaction. Most commerce systems still assume exactly that.

Agentic commerce breaks that assumption. When intent is generated outside the brand's owned environment, the handoff to transaction becomes a structural problem. Context doesn't transfer. Sessions don't persist. The consumer who asked an AI assistant for a recommendation and received one now faces the same friction-laden checkout process as someone who arrived with no prior intent at all.

Cart abandonment rates have remained stubbornly high for years. Baymard Institute research puts the average at 70%. That figure predates the agentic commerce era. As more purchase intent is generated through AI interfaces, and as the gap between that intent and a brand's transaction layer widens, the abandonment problem is likely to get structurally worse before it gets better.

What the current stack wasn't built to handle

The commerce infrastructure most enterprises operate today was assembled over two decades of incremental investment. Each layer added a capability: a search tool, a recommendation engine, a personalization layer, and a checkout system. Each was built to solve a specific problem within a human-initiated shopping journey.

None of it was built to receive intent from an AI agent.

When an AI system generates a purchase recommendation, it needs to do more than surface a product page. It needs to verify real-time inventory. It needs to apply pricing logic and promotional rules. It needs to respect brand policy around which products can be recommended together, which channels apply which discounts, and what the correct fulfillment path looks like for a given consumer. And it needs to do all of that without breaking the conversational context that made the recommendation possible in the first place.

Current commerce stacks can't do this reliably. The systems that hold the relevant data, inventory, pricing, order management, fulfillment, are not exposed in ways that AI agents can safely and accurately access. The result is a journey that starts with intelligence and ends with a broken experience: a link out to a product page, a generic checkout flow, and a consumer who arrived ready to buy and left without completing the transaction.

The conversion problem is an architecture problem

The industry has treated conversion optimization as a front-end problem for most of its history: better copy, cleaner checkout UX, fewer form fields, smarter retargeting. Those interventions were appropriate for the model they were built to serve.

The agentic commerce era introduces a different kind of conversion failure, one that front-end optimization cannot fix. When intent is generated externally, conversion depends on whether the back-end infrastructure can receive that intent, act on it accurately, and complete the transaction within the guardrails the brand has established. That is not a UX problem. It is an infrastructure problem.

Brands that are investing heavily in AI-powered discovery while leaving their execution layer unchanged are widening the gap between the promise AI makes on their behalf and the experience they can actually deliver. That gap has a cost, measured not just in lost transactions but in consumer trust that erodes each time the promise and the reality don't match.

Rezolve Ai commissioned research across 1,500 US consumers in January 2025 that found consumers who encounter friction immediately after an AI recommendation are significantly less likely to complete a purchase than those who encounter friction at the top of a traditional funnel. The implication is direct: AI raises the expectation bar at the moment of intent. Brands whose infrastructure cannot clear that bar are paying a conversion penalty they may not even know they're incurring.

What closing the gap requires

Closing the gap between AI-generated intent and completed transaction requires rethinking which layer of the commerce stack carries the most strategic weight in an agentic world. For most of the past decade, that weight sat with discovery and experience. The brands that invested most in search, personalization, and content won a disproportionate share.

In the agentic era, the weight shifts to execution. The brands that can reliably take AI-generated intent and turn it into a governed, accurate, brand-safe transaction will have a structural advantage over those whose infrastructure stalls at the handoff.

That is a different investment thesis than the industry has operated on. And most enterprise commerce roadmaps have not yet caught up to it.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

Mistral AI wants to build 1 gigawatt of European compute by 2030 β€” and lock in customers now.

Mistral AI wants to turn European AI sovereignty from a talking point into a product β€” one with a service-level agreement attached.

The French artificial intelligence company announced Tuesday a three-part expansion of its infrastructure business: regional inference endpoints that let customers choose whether their AI workloads run in Europe or the United States, a new "Priority Tier" backed by an uptime guarantee for mission-critical deployments, and a coalition of European enterprises making multi-year compute commitments that Mistral says will underwrite 200 megawatts of infrastructure across Europe by the end of 2027 β€” and a full gigawatt by the end of 2030.

In a move that may raise eyebrows among sovereignty purists, the company also said it will begin hosting third-party open models on its platform, starting with GLM-5.2 from Z.ai, the Chinese AI lab formerly known as Zhipu.

Taken together, the announcements mark a decisive shift in how Mistral positions itself. The company that built its reputation training open-weight language models is now selling something closer to critical infrastructure: assured capacity, regional control, and contractual reliability for enterprises and governments that want frontier AI without surrendering control over where it runs.

"When we spoke in June, the story was around how Mistral was building a full-stack AI offering," TimothΓ©e Lacroix, Mistral's co-founder and chief technology officer, told VentureBeat in an exclusive interview ahead of the announcement. "Today, the announcement is about strengthening one part of this infrastructure, which is the inference part."

That one part, it turns out, comes with a price tag measured in the tens of billions of dollars.

Inside Mistral's plan to build 1 gigawatt of European AI compute by 2030

The headline numbers deserve scrutiny, because they imply staggering capital requirements. Mistral currently operates less than 200 megawatts of capacity, according to the company. Details shared with VentureBeat show the near-term buildout resting on three sites: a 44-megawatt facility near Paris that became operational in the second quarter of this year, a 23-megawatt facility in Sweden built in partnership with EcoDataCenter using renewable energy and advanced cooling, and a 10-megawatt site in Les Ulis, France, that came online in the third quarter.

Getting from there to one gigawatt by 2030 is a different order of magnitude. Independent estimates suggest just how different: research firm Epoch AI calculates that a typical one-gigawatt AI data center requires roughly $38 billion in upfront capital expenditure, with servers and GPUs β€” not buildings or land β€” consuming the majority of the cost. Goldman Sachs Research pegs next-generation AI facilities at $15 million to $20 million per megawatt before accounting for the chips inside them.

Lacroix did not dispute the scale of the challenge. The investment required for a gigawatt of capacity "is a large investment that requires also a lot of scaling and revenue behind it," he said.

The urgency, in his telling, comes from a supply crunch that is about to get worse. "More and more, and especially around 2027 and 2028, we see that the demand for AI compute is exceeding what the market has to offer, especially in Europe," Lacroix said. McKinsey has estimated that meeting global AI demand could require $5.2 trillion in data-center capital expenditure by 2030 β€” and Europe, by most analyses, is starting from behind.

A company valued at a fraction of its American rivals cannot close that gap with venture capital alone. Which explains the most consequential β€” and most unusual β€” piece of Tuesday's announcement.

European Compute Units turn AI sovereignty into a five-year contract

Mistral is assembling what it calls an anchor group of enterprises whose long-term commitments will collectively finance infrastructure none of them could justify alone. Those commitments convert into "European Compute Units," or ECUs β€” a claim on Mistral-built capacity over multiple years that participants can spend on inference, training, model adaptation, or other AI workloads as their needs evolve.

If that structure sounds more like a power-purchase agreement than a cloud contract, that appears to be the point. Data-center financing increasingly resembles large infrastructure projects β€” gigawatts, substations, energy agreements β€” rather than traditional technology spending, and lenders want demand locked in before capital gets deployed. Mistral raised €830 million ($962 million) in debt earlier this year to fund its data center near Paris, TechCrunch reported in March, and pre-committed enterprise demand is exactly what makes that kind of financing repeatable at ten times the scale.

Lacroix was unusually direct about the mechanics. "The entire point of compute units is to have commitment," he said. "The goal is to have customers commit for around five years, or at least a long time." Asked what happens if a customer wants out early, he didn't soften the answer: "There is no getting out."

What makes a five-year, no-exit commitment palatable, he argued, is flexibility in how the capacity gets consumed. "Typically this can be spent on raw inference that you then feed through any other AI stack. It can be spent on raw compute as managed Kubernetes, and it can be spent at the very top with our full AI offering," he said. "My hope is that they will use it with our full-stack services and will love it."

The anchor group already includes some of Europe's industrial heavyweights. Amadeus CEO Luis Maroto said in a statement that "capacity, deployment control, and operating continuity become increasingly important for all enterprises." ASML chief Christophe Fouquet β€” whose company led Mistral's $13.4 billion (€11.7 billion) Series C last year β€” called building European AI capacity one of the few industrial endeavors that "will matter more to Europe's next generation," while Capgemini's Aiman Ezzat framed it as "a question of who shapes the future of European industry." CMA CGM chairman Rodolphe SaadΓ© said the shipping group's Mistral deployment is "already under way among thousands of employees."

Commitments of that duration only make sense, of course, if the sovereignty being purchased is real. On that question, Mistral's announcement contains an asterisk worth reading closely.

The fine print on sovereign AI: what data can still leave Europe

The centerpiece product is Mistral Regional Endpoints, now generally available, which let customers pin inference and its associated processing to Europe or the U.S. Alongside it, the new Priority Tier β€” in public preview β€” offers committed service levels, custom rate limits, and an uptime SLA for mission-critical workloads.

Mistral claims it is the only European AI lab offering both a choice of processing region and an SLA-backed service tier, and Lacroix said a third option is coming: an endpoint "that stays on Mistral-controlled infrastructure, so on Mistral compute" β€” for customers who want their inference not just in Europe, but off hyperscaler hardware entirely.

Then comes the fine print. Mistral's own materials note that in-region inference remains subject to "limited, safeguarded transfers" to sub-processors that may sit outside the chosen region. Pressed on what actually leaves Europe, Lacroix pointed to the connective tissue of modern AI applications: tool calls.

"There are some tool services, like some tool calls, that might be hosted in places where we don't fully control this," he said, citing web search as an example. "A few of our web-search providers might not all be in Europe, and in that case, we need to potentially gate that capability."

His answer to the compliance question β€” would this satisfy a European bank or a defense ministry? β€” was that gating is the feature, not the bug. Capabilities that cannot be sourced in-region can be switched off entirely, restricted to certain users or workspaces, or, given sufficient demand, rebuilt with European providers. "Any capabilities that we don't find a provider for in Europe β€” if it needs to be done in Europe, we'll find some way to implement it or find ways to address it," Lacroix said.

For enterprise buyers, that is a more honest framing than most sovereignty marketing offers: full regional control is available, but the moment an AI agent reaches out to the open web, sovereignty becomes a configuration decision rather than a default. The same pragmatism runs through the announcement's most surprising line item.

Why Europe's open source AI champion is hosting China's GLM-5.2

A French national champion β€” one that has partnered with the French army and positioned itself as Europe's answer to American AI dependence β€” hosting a Chinese lab's model invites an obvious question. Lacroix's answer was disarmingly matter-of-fact.

"It's a great model. Everyone loves it. It's open weight, so there was no good reason for us not to do it, really," he said, noting that Mistral's own stack is already built on open-source software like Kubernetes.

On security vetting, he argued that open weights fundamentally change the risk calculus. "The risks in taking a new model, at the layer of the weights, are β€” at least in my opinion β€” rather limited," Lacroix said. "We checked basically all of the safety and compliance evals that we have. We'll control that model, its outputs, and what it does the same way we do any of our models. We have the same inputs and outputs and monitoring capabilities over all of it."

The strategic logic is worth unpacking. By hosting third-party open models under European regional controls and the same SLAs as its own, Mistral is repositioning itself from model vendor to sovereign distribution layer β€” the trusted intermediary through which any open model, regardless of origin, can be consumed by a regulated European enterprise that could never call a Chinese API directly. It is the "model garden" playbook the hyperscalers run with Bedrock and Vertex, executed on European soil with European guarantees.

Customers appear to be reading it that way. "Mistral allows us to run open models under strict regional controls and service commitments, making it easy for us to maintain data residency and compliance requirements," Matan Griberg, CEO of AI software-engineering company Factory, said in a statement.

Lacroix stressed the move is not a retreat from frontier training: the model Mistral had in training as of June "is still training, and we're still very excited about it," he said. But openness to rivals' models signals where the company now believes its moat lies β€” not in any single model, but in the infrastructure underneath all of them. Which makes its relationship with the world's most powerful infrastructure company all the more interesting.

How the multibillion-dollar Microsoft deal funds Mistral's independence

Hovering over every sovereignty claim is Mistral's deepening relationship with Microsoft. In July, the two companies announced a multibillion-dollar expansion of their partnership under which Microsoft will rent capacity from Mistral's European data centers to serve its own cloud and AI demand, while adding Mistral Medium 3.5 and OCR 4 to Microsoft Foundry, bringing Medium 3.5 to Copilot Studio, and enabling Mistral models on Azure Local for disconnected, customer-controlled environments. Mistral CEO Arthur Mensch told The Wall Street Journal at the time that two-thirds of Mistral's customers already work with Microsoft.

How does a company selling independence from U.S. hyperscalers square taking one on as its largest tenant? Lacroix described Microsoft not as a patron but as an anchor customer that de-risks the buildout.

"It allows us to scale different parts of the business differently by building infrastructure with Microsoft as a customer," he said. "We can scale that team, we can scale our infrastructure, and make sure that we can then, on the side of it, also build for ourselves and for our customers." He compared the arrangement to the neocloud playbook β€” companies that built businesses supplying capacity to the hyperscalers themselves. "As that part of our business resembles that of neoclouds, we're following the same thing."

It is a genuinely clever inversion: rather than renting American infrastructure, Mistral is renting infrastructure to one of America's largest companies, using Microsoft's demand to finance capacity that also serves European sovereignty customers. But the independence has limits no contract can engineer away β€” the GPUs filling Mistral's European data centers come overwhelmingly from Nvidia and other American chipmakers, as SiliconANGLE noted in its coverage of the July deal.

Asked directly why a customer should choose Mistral over an EU region on AWS or Azure, Lacroix gave two answers. "The simplest possible answer is capacity. There is more demand than supply right now, and so it adds another option," he said. The second cuts closer to the pitch: "We are a European provider, and on the region that would be Mistral compute, we are fully independent. That's a truly differentiated offering than all of the hyperscalers or pure inference companies can provide."

The economics of open models: why agentic AI is pushing inference to the cloud

There has always been a tension at the heart of Mistral's business: its best-known models are free to download, and open models have historically been difficult to monetize through APIs. Asked how free weights fund a gigawatt buildout, Lacroix offered the clearest articulation yet of the company's thesis β€” that the economics of self-hosting are collapsing under the weight of the models themselves.

"When the models were smaller, and we were before the explosion of agentic AI, it was doable for enterprises to host their own β€” up to, let's say, 100-billion-parameter dense models β€” on their premises," he said. "More and more, with models going into the trillion or more parameters, with the current hardware, and with the increasing amount of tokens that need to be processed, it becomes harder."

His conclusion was blunt: "I don't see how, with the current trend of model size and growth of agentic tokens, we keep the full inference on-prem. To me, that is why we think we're going to monetize our cloud inference." Inference, he noted, is particularly well suited to the cloud because it "does not need to hold any data" and can be encrypted in transit.

In other words: open weights get Mistral into the enterprise, and the physics of trillion-parameter agentic workloads brings the inference β€” and the revenue β€” back to Mistral's data centers. The thesis will get an expensive test. Mistral has raised roughly $4 billion to date, according to PitchBook data β€” a fraction of the war chests assembled by OpenAI and Anthropic β€” and Bloomberg reported in June that the company is in talks to raise about €3 billion at a roughly €20 billion valuation, nearly double its Series C mark. The revenue behind the buildout will have to come from exactly the enterprises Tuesday's announcement is courting.

And Europe, in Mistral's telling, is only the first market for what it is selling. Asked whether the framework could be replicated in the Middle East, Asia, or anywhere else anxious about AI dependence, Lacroix didn't hedge: "It's completely right. We're starting this in Europe because it's also an easier part of the world for us to scale into, especially in the infrastructure. But we definitely want to extend this, depending on customer demand." Every layer of the stack, he said, "can be controlled, changed, replaced depending on where we operate and what the requirements are β€” that's pretty much where we excel."

That is the wager underneath the SLAs, the compute units, and the Chinese model flying a European flag: in a world where the U.S. and China dominate frontier AI, the durable business is selling everyone else control. To fund it, Mistral is asking Europe's largest enterprises to sign five-year contracts with no exit β€” while making a bigger, longer commitment of its own. A gigawatt, after all, is a promise measured in decades. For Mistral, too, there is no getting out.

Nvidia's Switchyard router reshuffles AI models mid-task, cutting task costs to a third in its own tests

11 August 2026 at 13:00

Enterprises running always-on AI agents keep hitting the same tradeoff. Send every task to a frontier model and the bill climbs fast. Build custom routing logic to send easy tasks to cheaper models and that becomes its own engineering project, one that has to be maintained every time a workflow changes.

Nvidia is proposing a fix that touches both ends of that problem at once. The company is out on Tuesday with Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for high-volume, specialized agent tasks, alongside NeMo Switchyard, an open-source library that routes each step of an agent workflow to whichever model fits it best.

The headline numbers: According to Nvidia, Lightning delivers up to 4x faster output than comparable models in its class, completing agentic tasks roughly 30% faster than Qwen3.6-35B at matching accuracy. Paired through Switchyard, Nvidia says the combination holds frontier-level task completion while cutting benchmark costs to roughly a third of running Opus 4.8 alone.

The timing puts Nvidia in the middle of the busiest open-weight stretch the industry has seen in months. Alibaba, Moonshot, Zhipu and DeepSeek have all shipped competitive open models out of China since the spring, several landing at or near frontier performance while undercutting US labs on size or price. Meta added to that pressure by releasing its own 30-billion-parameter open agentic model, Muse Glimmer. Open weights have gone from a differentiator to table stakes in a matter of months, and Nvidia's release lands squarely inside that shift rather than ahead of it.

The pairing is the point. A model alone doesn't solve the cost problem, and a router alone has nothing efficient to route to. Nvidia is betting that open source, applied at both the model layer and the routing layer, is what actually moves the cost needle on agentic AI, not a single cheaper model and not a smarter router bolted onto someone else's stack.

Switchyard's real rivals aren't other open models β€” they're Not Diamond, which already powers OpenRouter's Auto mode, and RouteLLM, the open-source framework from UC Berkeley and LMSYS. Neither ships its own model. Nvidia's bet is that owning both sides of the decision, under one open license, is what a router-only or model-only competitor can't match.

"That is the power of a system of models, matching the right model to each step of the workflow," Kari Briski, vice president of generative AI at Nvidia, said in a briefing.

How the router actually changes the workflow

Model routing isn't a new category. OpenRouter, LiteLLM and a handful of standalone routing startups already let developers point traffic across multiple providers. Switchyard plugs into several of them rather than replacing them outright.

The core problem Switchyard solves is that the right model changes as an agent moves through a task. An agent's state shifts as tools return results, errors show up, or a step turns out to be routine rather than complex, and a fixed model choice can't adapt to any of that.

Briski described routing strategies that respond to that shifting state rather than a static task category.

"It has many types of routing strategies," Briski said. "You can have a random router, which is not that great, or you can have an agent state route or a classifier route. Depending on your routing strategy, it wants to choose the best model. In some cases you want to go with a model like Lightning for really efficient tasks, and the router will actually choose Lightning if it's set up in your pool of models."

Cost enters the routing decision directly, not as an afterthought. In response to a question from VentureBeat, Briski said Switchyard can evaluate model verbosity, meaning how many tokens a given model tends to produce for a task, and use that prediction to steer work toward the cheaper option before the call is made.

The part that keeps this from becoming its own integration project is where Switchyard sits. Nvidia split its partners into two groups: agent frameworks that call Switchyard directly, including Cognition, LangChain and Nous Research, and LLM gateways that have built Switchyard support into their own products, including Kong, LiteLLM and OpenRouter. Kong ships Switchyard natively inside Kong AI Gateway. Briski pointed to that same list of gateway partners when describing how the library fits into the existing routing ecosystem.

"We are an ecosystem lover, and we want to make sure that we are integrated," Briski said. "We've partnered with OpenRouter, LiteLLM and Kong, and they've already integrated our routing algorithm, so you can pick it up right where you're already using the best tools."

Nvidia shared results from nine companies testing Switchyard, several with specific figures attached. LangChain reported a 74% cost reduction across 145 multi-turn Deep Agents tasks by routing just 7% of calls to a frontier model, at a 6% accuracy tradeoff. Ramp said it matched a frontier model's performance on Ramp SWE-Bench while cutting costs 58% and runtime 33%. Cognition integrated Switchyard's staged router into Devin Desktop for internal use and reported near-frontier performance on FrontierCode Main while cutting mean cost 28% relative to routing everything to a single frontier model.

Lightning's architecture and performance gains

Nemotron 3.5 Lightning is a standalone open model in its own right, built for high-volume, specialized agent tasks rather than general-purpose use.

It extends the hybrid Mamba-Transformer, latent mixture-of-experts architecture Nvidia introduced with the Nemotron 3 family in December 2025, the same line behind Nemotron 3 Super, which Nvidia uses as Lightning's own baseline in its post-training comparisons. Positioned within a routing setup like Switchyard, it's built to sit at the fast, cheap end of the decision rather than the frontier end, but it runs and ships independent of any router.

According to the Artificial Analysis Intelligence Index, a general capability benchmark spanning nine evaluations, Lightning scores 24, tied with gpt-oss-120b and behind Nemotron 3 Super, Gemma 4 31B, Claude 4.5 Haiku and Mistral Medium 3.5, all at 30. Lightning isn't a general-intelligence leader in its size class, and Nvidia isn't claiming it is.

The actual claim is narrower: according to PinchBench data supplied by Nvidia, Lightning matches Qwen3.6-35B's accuracy roughly 30% faster and beats Gemma 4 26B's accuracy at a similar completion time on PinchBench, a real-world agent task benchmark spanning coding, research and file management. That's a speed-to-accuracy tradeoff, not a capability win.

Post-training is where Nvidia says the bigger gains show up. The company shared before-and-after figures from four early-access partners: CrowdStrike's malicious-content recall against a Nemotron 3 Super baseline, CodeRabbit's coding router against a GPT 5.4 Nano baseline, Harvey and Trajectory's legal task completion against an Opus 4.6 baseline, and Lila Sciences' energy simulation work against an Opus 4.8 baseline. CodeRabbit's case is the most specific: Nvidia says the standard NeMo Auto model recipe, trained for one epoch, built into a working router agent for $85 in about two hours.

What this means for enterprises

There is no shortage of competitive offerings in the growing market for open models. The new Nemotron Lightning release will be yet another option for organizations to consider.

On the model side, Lightning's own benchmark chart picks Qwen3.6-35B as its direct comparison point. Asked by VentureBeat directly how Lightning compares to Chinese models more broadly, Briski didn't offer a head-to-head benchmark, pointing instead to openness and customizability as the differentiator.

"Our value proposition is not just open and it's very customizable," Briski said.

For enterprises building agentic infrastructure, three trends stand out:

The routing decision is becoming dynamic instead of static. Enterprises that built agent pipelines around a single default model are being pushed toward per-step routing based on live signals like agent state and token cost, not a fixed assignment set at design time.

Open source is now a cost lever at two layers, not one. Pairing an open model with an open router a vendor controls end to end is a newer argument than cheaper weights alone, and worth watching for whether other labs follow the same pattern.

The competitive question shifts from best model to best system. As routing libraries mature, the differentiator moves from which model an enterprise defaults to, toward how well its routing layer matches models to tasks in production, a harder thing to benchmark and a harder thing to market.

Your AI agent may be ready. Your sales motion probably isn’t.

11 August 2026 at 13:00

Presented by Salesforce


Interested buyers don't generate revenue. Live customers do. That's the lesson I keep drawing from watching hundreds of ISV partnerships navigate the agent economy over the last 18 months.

The companies pulling ahead aren't winning on features. They're winning because customers can move from discovery to live deployment in hours, while competitors are still negotiating contracts, clearing tax reviews, and waiting on provisioning.

That gap between a buyer who says β€œyes” and a customer who is actually using the product is where too many deals lose momentum. Urgency fades. Champions move on. Competitors get another opening.

Gutenburg saw that gap firsthand. Healthcare organizations valued its product, but sales cycles stretched 30 to 45 days. With custom pricing via AgentExchange, the company closed an urgent healthcare deal in just 48 hours.

Not 48 days. 48 hours.

The final contract phase alone dropped from 4 hours to 4 minutes. A 60x improvement.

I see this pattern across the ISV ecosystem. Building agents is getting faster. Getting buyers live before urgency fades is becoming the constraint. In a market moving this quickly, that can matter as much as the agent itself.

It’s like building a bullet train and selling tickets by fax. The product is built for speed. The transaction is not.

Distribution beats product in crowded markets

Nearly every software company is pouring resources into agent development. Far fewer are rethinking the path from discovery to deployment. Manual contracts, custom invoicing, tax reviews, provisioning delays, these are the handoffs that turn a 48-hour deal into a 45-day cycle.

That friction is now a competitive disadvantage, because the buying process is changing faster than most back offices are. Gartner predicts that by 2028, 90% of B2B purchases will be guided by AI agents.

That does not mean humans disappear from enterprise buying. It means the discovery and evaluation process changes. Buyers will increasingly use AI to identify, compare, and narrow solutions.

If your agent is not discoverable where that evaluation is happening, you may never make the shortlist.

A better agent can still lose to one that's easier to buy.

Domain expertise matters. Workflow depth matters. Proprietary data matters. Customer context matters.

But enterprise categories are getting crowded fast. In crowded markets, the best product does not always win. The product that is easiest to discover, buy, deploy, and scale often has the advantage.

As agent-guided buying takes hold, the first evaluation may happen before a demo is scheduled or a sales rep is in the room.

AI agents will increasingly scan marketplaces, compare solutions, and help narrow purchase decisions in the time it used to take to schedule a discovery meeting.

Companies that figure out marketplace distribution now will own their categories.

That is the problem AgentExchange was built to address. It’s a single destination for apps, agents, and integrations that extend and connect to Salesforce and Slack, helping customers get more from their platform investments.

But discovery is only the first step. The bigger question is what happens after the buyer says β€œyes”.

β€œYes” doesn't mean live

Enterprise software teams spend enormous energy getting to "yes." But in many deals, that is where the operational work begins.

Between β€œyes” and β€œlive,” the back office can generate a chain of handoffs: contracting, invoicing, tax calculation, licensing, provisioning, fulfillment, payment, and finance reconciliation. Every handoff delays activation for the customer and delays recognized revenue for you.

For AI agents, that back-office drag is becoming a front-office problem.

AgentExchange brings discovery, commerce, and activation together, helping partners manage custom pricing, billing, licensing, provisioning, and fulfillment through one connected experience.

"AgentExchange removes the traditional procurement friction that slows deals. Customers can now discover, purchase, and deploy PandaDoc directly through their existing Salesforce contract, turning what used to be a multi-week process into a same-day activation." Keith Rabkin, CEO at PandaDoc

What closing in 48 hours actually looks like

Gutenburg’s 30-45 day cycles were eaten up by contract logistics. Sales moved faster than their back office.

Using custom pricing and automated transaction capabilities through AgentExchange, they streamlined contracting, tax calculation, provisioning, and other steps between buyer interest and activation.

When a healthcare organization needed a tool to help them create documents aligned to the Americans with Disabilities Act and accessibility requirements, Gutenburg closed in 48 hours from first contact.

The 48-hour close is the differentiator. It is what efficient growth actually looks like in practice. Revenue scales without scaling headcount. Pipeline coverage improves because you are discoverable everywhere. Net recurring revenue increases because customers expand through the same frictionless channel.

"AgentExchange condenses contracting and tax calculations into a 10-minute process with improved accuracy," said Zamial Jones, VP of Customer Success at Gutenburg. "For partners spending hours on these tasks for every deal, that's transformational."

The window is closing faster than you think

The app economy took a decade to mature.

The agent economy won't.

The ISV partners I've watched pull ahead aren't the ones with the most sophisticated agents. They're the ones who treated distribution as a product problem β€” resourced, measured, and iterated β€” before the category consolidated around them. The ones still treating go-to-market as a post-launch consideration are consistently 6 to 12 months behind.

You can spend the next two quarters perfecting your agent's reasoning capabilities. Or you can spend them making sure customers can actually buy it.


Salesforce is investing in the next generation of companies creating agents with $50 million through the AgentExchange Builders Initiativeβ€”capital, engineering support, co-marketing, and co-sell programs. Companies that move now will define what enterprise AI distribution looks like for the next decade. Learn more here.

Lisa Eisenberg is SVP of ISV Partnerships at Salesforce.


Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

LTX-2.5 can generate a 10-second AI video from an image in just 6.8 seconds on Nvidia superchips β€” and it's open weights

LTX, the open world model company spun out of Lightricks, today released LTX-2.5, the newest version of its open-weights video and "world" model and it arrives natively integrated into ComfyUI, the node-based workflow tool that has become the de facto prototyping environment for open generative media, through a strategic day-one launch partnership between the two companies.

The model is available now as open weights on Hugging Face, inside ComfyUI, and through the LTX API for teams that want managed generation. It is free to use for organizations under $10 million in annual recurring revenue; larger companies negotiate a license. LTX says its models have passed 33 million downloads, making the LTX family the most-used "open world" model line on the market.

Ahead of the launch, VentureBeat spoke exclusively with LTX co-founder and CEO Zeev Farbman and ComfyUI co-founder and CEO Yoland Yan about the release, the partnership, and why both companies are betting that open weights β€” not closed APIs β€” will win the video and world model market.

"We're trying to maintain the same efficiency and the inference speed that we're known for, but constantly pushing the quality up," Farbman said. "We are introducing many cool things in this release: multi-shot support, a diffusion decoder for better quality, new conditioning modes, better support for autoregressive models that are critical for real-time use cases and robotics."

What's new in LTX-2.5

According to the company's announcement, LTX-2.5 rebuilds nearly every stage of the generation pipeline rather than bolting new capabilities onto an older core. The headline changes:

  • A new diffusion video decoder that reduces visual artifacts in high-motion footage and reconstructs fine detail like text and faces, while preserving LTX's high compression ratio.

  • Native multishot generation that renders a full sequence as a single output, holding character, scene, and voice consistent across cuts rather than stitching individually generated shots together.

  • A custom Gemma 4 language backbone and dedicated prompt enhancer for more accurate handling of complex, multi-subject prompts.

  • A pretrained checkpoint tuned for physical AI and robotics giving teams a base to fine-tune on domain data that looks nothing like cinematic video.

  • A substantially improved distilled model that delivers near-full-model quality at lower cost and faster inference, and, through an optimization effort with NVIDIA, runs locally on NVIDIA RTX GPUs with reduced memory requirements.

The company claims roughly one-eighth the cost and one-seventh the render time of comparable models, with output that runs on hardware ranging from data center GPUs down to a Mac.

Checked against published rates, the cost multiple doesn't survive contact with the models that publish pricing.

LTX-2.5 generates 720p video with audio at $0.09 per second on its Fast tier, putting a 10-second clip at $0.90 β€” genuinely cheap, but about one-quarter the cost of full Veo 3.1 ($4.00), half of FLUX 3 Video ($1.70) and HappyHorse 1.0 (~$1.82), and only 10% under Google's budget tiers, Veo 3.1 Fast and Gemini Omni Flash ($1.00 each), while Veo 3.1 Lite ($0.50) is actually cheaper.

Nothing in the published field costs eight times LTX's rate; if the one-eighth figure holds anywhere, it would be against premium models like Kling 3.0 Pro or Seedance 2.5 that don't publish comparable per-second pricing β€” or against self-hosting the open weights, where the marginal cost is whatever your GPU costs to run.

The render-time multiple is better supported, at least by LTX's own end-to-end measurements: 6.8 seconds for a 10-second clip against 52 seconds for the fastest rival API (Gemini Omni Flash) is roughly one-seventh β€” though that figure comes from self-hosting on two GB200 superchips, and through LTX's own managed API the same job took 23.7 seconds, cutting the advantage to about half.

Here is how the published rates compare, normalized to the cost of a finished 10-second 720p clip with synchronized audio β€” the configuration LTX and Black Forest Labs have both used for their own evaluations:

Rank

Model

Per second

Per 10-second clip

Unique differentiator

Notes

1

Veo 3.1 Lite (Google)

$0.05

$0.50

The category's price floor β€” cheapest published rate anywhere

No 4K, no clip extension

2

LTX-2.5 Fast (Lightricks)

$0.09

$0.90

Only open-weights model in the field β€” self-host free under $10M ARR, fine-tuning permitted

Scales to 4K at $0.30/sec; up to 20s single generation at 24/25 fps

3

Veo 3.1 Fast (Google)

$0.10

$1.00

Cheapest closed-API path to 4K ($0.30/sec)

Budget tier of the Veo line

3

Gemini Omni Flash (Google)

$0.10

$1.00

Independently measured quality leader β€” tops both Artificial Analysis text-to-video arenas as of Aug 2026

720p only; 10-second maximum; best iteration tooling

5

LTX-2.5 Pro (Lightricks)

$0.12

$1.20

Quality-tuned tier of the only open-weights family β€” prompt adherence, faces, typography (vendor-described)

Tops out at 1080p and 10 seconds

6

FLUX 3 Video (Black Forest Labs)

$0.17

$1.70

First to ship 20-second single-generation clips with audio (July 2026) β€” a ceiling since matched by LTX-2.5 Fast

HD band; audio included; Draft tier at $0.06/sec ($0.60/clip, HD only)

7

HappyHorse 1.0 (Alibaba)

~$0.182

~$1.82

Arena quality leader at launch (April 2026), since overtaken; "open source" claims never matched by verified downloadable weights

Third-party reseller rate; audio included at no extra charge

8

Veo 3.1 (Google)

$0.40

$4.00

Only model supporting clip extension beyond a single generation

Premium tier; 8x Veo 3.1 Lite; 1080p at no premium over 720p

Sources: LTX API pricing documentation; bfl.ai/pricing; Google AI for Developers model pricing; HappyHorse reseller rates via third-party API platforms; Artificial Analysis text-to-video arena leaderboards. All rates verified August 11, 2026, and subject to change.

How fast and how good LTX says it is

The most eye-catching number in LTX's launch materials is speed: the company says LTX-2.5 generates a 10-second, 720p image-to-video clip in 6.8 seconds faster than real time.

The caveat is the hardware behind it. That figure was measured self-hosted on two of NVIDIA's top-end GB200 chips at steady state, a configuration far beyond what most teams have racked; the same job through LTX's own managed API took 23.7 seconds, albeit rendered at the higher 1080p resolution (the API has no 720p tier).

By the company's end-to-end measurements of competing APIs on the same task, Google's Gemini Omni Flash came in at 52 seconds, xAI's Grok 1.5 at 63 seconds, Google's Veo 3.1 at 70 seconds (for an 8-second clip), MiniMax H3 at 180 seconds, ByteDance's Seedance 2.5 at 317 seconds, and Kuaishou's Kling 3.0 Pro at 398 seconds.

On quality, LTX shared results from blind, side-by-side human preference tests, in which evaluators voted on videos generated from the same prompt without knowing which model produced which.

LTX-2.5 recorded a 67% win rate, narrowly ahead of Seedance 2.5 at 65%, with Gemini Omni Flash at 55%, MiniMax H3 at 50%, Seedance 2.0 at 44%, Wan 2.6 at 42%, and FLUX 3 at 28%.

All of these figures are vendor-reported measured or commissioned by LTX itself, not independently verified and the company labels the preference results preliminary, noting it expects them "to evolve as evaluation expands." They are directional claims a buyer should test against their own workloads rather than settled rankings.

The independent benchmark that does exist cuts the other way for now: as of this month, Gemini Omni Flash β€” which LTX's commissioned tests place 12 points behind its own model β€” leads both of Artificial Analysis' text-to-video arena leaderboards, and the arena does not yet score LTX-2.5 at all. Until it does, the 67% figure remains untested on neutral ground.

The launch materials also lean on deployment terms rather than raw performance: LTX-2.5 runs on any GPU with a minimum of 16GB of VRAM, deploys on-premises, at the edge, or via API, carries no visible watermark on output β€” though the license requires users to disclose that content is machine-generated and forbids removing any embedded provenance or "latent disclosure" features (more on this below) β€” and can be fine-tuned on a customer's own data and IP flexibility the company contrasts with closed API-only rivals and with open-licensed competitors whose weights are unavailable in the U.S. and Europe or whose licenses restrict fine-tuning.

Betting against the API business model

For Farbman, the release is another installment in a strategy that began as a reaction to the industry's consolidation around closed models.

"We started with our own models out of necessity, because around the time that Sora came out, we realized that all the big guys are trying to close their models, and working through APIs just doesn't work for many businesses, including the kind of stuff that we wanted to build," he said.

The technical argument, he explained, is that video and world models have a fundamentally wider "surface area" of use cases than language models.

"With LLMs, the surface area of the API is pretty narrow, we're typically asking some kind of question, passing words and getting words back," Farbman said. "With video models, world models, there are so many different use cases that require people to get access to the weights and create flows that really work for them."

He was blunt that the openness is not charity. "We're definitely not doing this as philanthropy," he said. "Our answer is open weights with licenses that allow individuals and companies below a certain amount of revenue to use the model for free, and once they're successful, to come up with some kind of licensing agreement with us."

"We're trying to build a model that builders can confidently build upon," he added. "We're coming and saying: guys, open weights is not some kind of one-time philanthropic fluke for us. It's the strategy. We believe this is the right way to serve these models, and we're going to keep doing that."

What the license actually says

"Open weights" and "open source" part ways in the fine print. LTX-2.5 ships under the LTX-2.x Community License, a custom agreement that would not qualify as open source under the Open Source Initiative's definition: it discriminates by revenue and by field of use, both disqualifying restrictions.

The headline mechanic works as advertised β€” organizations are free to use, modify, self-host, and even sublicense the model, with the $10 million annual revenue threshold (measured across all affiliates and subsidiaries, so a small subsidiary of a large parent doesn't slip under it) triggering the paid license.

Notably, even companies above the line can download and evaluate the model free in non-production environments β€” the license effectively codifies the prototype-in-ComfyUI-then-license funnel Farbman describes. It also gives that funnel teeth: unauthorized commercial use obligates the violator to pay back-fees at LTX's standard rates, due within 30 days of written demand.

The stickiest provisions concern what counts as a "derivative." The definition sweeps in not just fine-tuned checkpoints and LoRA adapters but distillations and any model trained on LTX-2.5's outputs or synthetic data β€” meaning a company that generates training clips with LTX-2.5 and uses them to train its own unrelated model has, by the license's terms, created a derivative locked to the same agreement.

All derivatives must be redistributed under the same license, a fine-tune transferred to a $10 million-plus company triggers that company's own paid-license obligation regardless of who built it, and commercial users are barred outright from using the model to train or improve any competing AI system. A separate clause prohibits deploying LTX-2.5 in any product that competes with Lightricks' own offerings without a negotiated license.

There are also control provisions unusual for a self-hosted model. Lightricks claims no rights in generated output, but the license requires users to disclose that content is machine-generated, forbids removing or circumventing any watermarking, content-provenance, or "latent disclosure" features embedded in the model, and reserves Lightricks' right to restrict usage "remotely or otherwise" and to push updates β€” with immediate license revocation as the penalty for disabling disclosure features.

The license also declares Lightricks' intent that LTX-2.5 be treated as a "free and open-source general purpose AI model" under Article 53(2) of the EU AI Act, a derogation that lightens the company's own regulatory obligations β€” a classification legal observers may contest precisely because of the revenue threshold and use restrictions in this same document. And one restriction bears directly on the physical-AI pitch: military, warfare, and weapons-development uses are banned entirely, so the robotics checkpoint is off-limits to the defense sector without separate terms.

From Facetune to world models and the node graph that became a standard

LTX grew out of Lightricks, the Jerusalem-headquartered company best known for consumer creative apps including Facetune and Videoleap. Bootstrapped and profitable, Lightricks pivoted to foundation models in 2022, launched its LTX Studio filmmaking platform in early 2024, and released its first open-weights LTX Video model (LTXV) in November 2024, following it with a 13-billion-parameter version in May 2025. Farbman co-founded the company alongside CTO Yaron Inger and CMO Nir Pochter, and the LTX brand now fronts its world model business, with offices in New York, London, and Chicago.

ComfyUI began in January 2023 as an open-source side project by a pseudonymous developer known as "comfyanonymous," who built a node-based graphical interface for Stable Diffusion that let users chain models and processing steps into repeatable visual workflows. It has since become one of the fastest-growing open-source projects in generative media the standard environment where new image and video models are tested, combined, and pushed into production and is now backed by a company, Comfy Org, which raised $17 million to keep developing the tool. Yan, a co-founder, serves as its CEO.

Why ComfyUI is the front door for enterprise adoption

For readers wondering why a model company and a tooling company are launching arm-in-arm, Farbman's answer was unusually candid: ComfyUI is where LTX's paying customers come from.

"A whole lot of our customers are starting their journey with Comfy," he said. "It's already this prototyping system that's extremely popular in the industry, and a lot of the potential customers are coming to us after they already figured out the flow inside Comfy. It's already working, so for us it's a no-brainer that we have to provide zero-day support for the Comfy integration, because it's basically our customer acquisition channel."

Yan described ComfyUI's role as the connective layer of the open ecosystem. "Comfy at the core is sitting as a layer on top, giving people accessibility to the open-weight models that people can inference on their local machine, or tap into closed models as well through our partner node system," he said. "In the end, [they] combine everything together into a workflow that empowers various things, from the creative side all the way to data pipeline and robotics type of scenarios."

That flywheel, Yan argued, is what sustains open models commercially: "We help promote and push these models into the world... people do all sorts of workflow and model innovation on top of it, and that further propagates these models into studios or robotics labs. Those companies would end up acquiring licenses and then contribute a part of the value gained back to LTX and the rest of the ecosystem."

What enterprises should know

Both executives pushed back on the assumption that a video model is only for generating videos. Farbman rattled off a list of enterprise deployments that have little to do with cinematic clips.

"We have hardware customers that are trying to figure out how to do computational photography with diffusion models, for example, taking a stream of raw pixels that are coming from the sensors, which is typically very noisy, and trying to figure out how to reduce noise there," he said. "Or think about the production studios that are trying to figure out how to do VFX, how to do water simulation, how to turn day into night. Or think about animation studios: they're trying to figure out how to streamline their pipeline, where animators are creating keyframes and then the system uses them as interpolation."

For enterprises weighing where to start, the recommended path is the one their own employees have probably already taken. "A lot of enterprises have already adopted Comfy, and I think many others will follow," Farbman said. "It gives this right level of structure, where you can tweak things a lot, but it still abstracts a lot of things away... Enterprises are typically reaching out after people internally have already played with the model, played with Comfy."

Yan described a consistent two-track pattern among studios and companies already running LTX and other open models in production. "They have their research, or R&D, creative pipeline, anything goes," he said. "Once in a while, some of these pipelines get good enough that they graduate into some kind of production environment. And somewhere along the line, the enterprise conversation gets started. On our end, it's more around tooling, and on the LTX side, it's more around the licensing."

Because the weights are open, that entire experimentation phase can happen on a company's own hardware, with no per-generation billing and no data or IP leaving its systems, a meaningful distinction for enterprises with sensitive footage, proprietary characters, or regulated data. The commercial trigger only arrives with scale: organizations above $10 million in ARR need a license.

Yan framed the stakes for slower-moving companies in starker terms. "This is a trend that is just fundamentally going to disrupt the entire creative industry," he said. "Studios are heavily trying to figure out what is the roadmap and how do we get ahead, sometimes not even get ahead, just how do we avoid falling behind the AI adoption wave."

Developers, real-time apps, and the edge

For software developers, the release leans into a growing real-time story. Alongside ComfyUI, LTX named two other launch partners: Asteria, the AI film studio producing original film and video on LTX, and Reactor, a developer platform that runs LTX-2.5 on low-latency inference infrastructure to power interactive avatars, live worlds, and real-time robotics workloads, so developers can build production-grade real-time experiences without standing up that infrastructure themselves.

Yan pointed to a viral example of what open weights plus low latency makes possible: Flipbook, an interactive experience that spread on Reddit in which an entire clickable world is generated on the fly. "Everything people see on that interface is generated using an LTX model, live-streamed," he said. "It's an environment, or a world, where anywhere you click, it just generates a brand-new interaction... That type of experience and experimentation wouldn't exist without an open-weight model, without LTX's type of performance."

Farbman said efficiency at the edge is a deliberate design target, not a side effect. "For us, it's very important to create an extremely efficient model that people can run on edge devices, both on consumer hardware and close to the edge with physical AI," he said, while acknowledging the relentless pace of the field: "These days, it's almost hard to take a vacation. Things are progressing so quickly that while you're releasing one model, you're already deeply into training another one, and new papers are coming on a daily basis."

Filmmakers: virtual production now, easier slopes later

For professional filmmakers and studios, Yan sees real-time world models changing the shape of production itself, collapsing the gap between shooting and post. "These days you see real-time models, or world models, getting adopted in studios as part of what's called virtual production, meaning you can shoot and then immediately get close to what the post-production result looks like," he said. "You give a much better experience to the producer or director to say, 'okay, this is what I want,' or 'this is not what I want let me actually reiterate.' Whereas before, the entire Hollywood pipeline is, in my opinion, a giant mess where it has to constantly go between multiple departments."

He also cautioned against reading head-to-head model comparisons too literally, given how differently models specialize across animation, photorealism, gaming, 3D, and robotics. "Various models have simply different characteristics," he said. "It's like comparing Michael Phelps with, I don't know, Michael Jordan. It's not really a comparison of who's a better athlete, there are just different specialties here."

As for amateur and indie creators intimidated by ComfyUI's famously steep learning curve, Yan was direct that the tool will meet them partway, but only partway.

"It's kind of like skiing," he said. "There are easy slopes that you can go down using Comfy, and hopefully we can create more and more of these easy slopes overall. But we'll never sacrifice the existence of the double-black-diamond type of lanes, because the real technical, professional creatives actually need and couldn't live without that type of core power. That's actually our core differentiator compared to a mobile-app type of creative tool."

LTX-2.5 is available today on Hugging Face, natively in ComfyUI, and through the LTX API.

Updated several hours after publication with additional details from LTX's public blog post and API pricing page.

OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks

Earlier today, OpenAI launched GPT-5.6-Cyber, a specialized model designed to perform advanced vulnerability research and exploit development for approved defenders β€” including categories of work that its general-purpose models will often refuse.

GPT-5.6-Cyber is a fine-tuned version of OpenAI's most advanced general model, GPT-5.6 Sol, unveiled back in June, but trained specifically to improve performance on advanced cybersecurity tasks, including finding zero-day vulnerabilities and developing exploit chains.

Crucially, OpenAI also trained it to reduce refusals on some higher-risk, "dual-use" cybersecurity requests β€” that is, requests that could be used for legitimate defensive or malicious offensive purposes.

Indeed, on an internal OpenAI benchmark called Advanced Cybersecurity Completion Rate β€” which the company says in its launch blog post measures tasks involving exploit-chain development, authentication bypass, privilege escalation, and other advanced cybersecurity scenarios β€” GPT-5.6-Cyber completed 95% compared to just 57.3% from its immediate predecessor model GPT-5.5-Cyber, and just 1.5% with the normal GPT-5.6 Sol model and all its safeguards applied.

OpenAI researcher Eric Wallace posted on X, describing GPT-5.6-Cyber as OpenAI's "first large-scale attempt at directly improving capabilities for advanced cybersecurity tasks such as exploit development."

Pricing and availability

Unfortunately for enterprises, GPT-5.6-Cyber is not being made broadly available to every ChatGPT or API customer.

To get access, an organization has to be accepted into the newly created tier of OpenAI’s Daybreak cybersecurity program, called Daybreak Red β€” also announced today, which gives access to dedicated cybersecurity models like GPT-5.6-Cyber

Another new tier, Daybreak Blue, gives a wider swath of enterprises access to general models like GPT-5.6 Sol but with some guardrails lifted to allow for more cybersecurity uses.

OpenAI’s documents list pricing for GPT-5.6-Cyber at $12.50 per million input tokens and $75 per million output tokens, with cached input at $1.25 per million tokens.

That makes it more expensive than GPT-5.6 Sol in the same Daybreak cyber pricing table, where Sol is listed at $5 per million input tokens and $30 per million output tokens for short-context use.

OpenAI does not list long-context pricing for GPT-5.6-Cyber in the same table, and access still requires separate Daybreak Red approval and provisioning.

Red vs. Blue: OpenAI's new Daybreak tiers and how to qualify for them

Daybreak Red is for approved security teams doing advanced, authorized cyber work β€” the kind of work that can look risky out of context, even when it is being done for defensive reasons. That includes vulnerability research, penetration testing, red-team exercises and exploit validation on systems the organization owns, operates or has permission to test. In other words, OpenAI is saying GPT-5.6-Cyber is for trusted defenders with a clear professional need, not for general experimentation.

Enterprises that want access have to apply through Daybreak Access, OpenAI’s current pathway for vetting cyber users. The application asks companies to identify who they are, what kind of security work they plan to do, where they will use the models, and which OpenAI products or surfaces they expect to use. Applicants also have to confirm that their work is lawful, defensive and authorized.

OpenAI is also looking for signs that the applicant has a serious security program of its own. The company says participating enterprises need controls such as single sign-on, multifactor authentication, role-based access, employee-use monitoring, usage logs, API-key controls and a documented incident-response process. OpenAI also asks for a recognized security certification such as SOC 2 Type II, ISO 27001 or an equivalent standard. Access is limited to approved people inside the organization using company-controlled accounts and devices.

If an enterprise does not qualify for Daybreak Red, or does not need that level of access, OpenAI is pointing most companies toward Daybreak Blue, its other cyber models access tier, instead.

Blue is the broader tier for approved defenders. It does not provide GPT-5.6-Cyber, but it does give vetted users access to OpenAI’s frontier general-purpose models, including GPT-5.6 Sol, with safeguards adjusted for legitimate defensive work.

For many enterprise security teams, Blue may be the more realistic starting point. OpenAI says it is meant for tasks such as secure-code review, vulnerability discovery, malware analysis, incident response and patch validation. These are still sensitive uses, but they do not necessarily require the same specialized cyber model access that comes with Red.

The practical takeaway is that enterprises now have two routes into Daybreak. Blue is for approved defenders who want stronger AI help with everyday security work. Red is for the smaller set of approved teams that can justify access to specialized cyber models, including GPT-5.6-Cyber. Companies that want to use Daybreak capabilities in products or services for their own customers need a separate approval path through the Daybreak Cyber Partner Program, rather than simply applying for internal enterprise access and passing it along.

How OpenAI got here: from Trusted Access to Daybreak

OpenAI has supported defenders through its Cybersecurity Grant Program since 2023 β€” later expanded to $10 million β€” and began building cyber-specific safeguards into its model deployments starting with GPT-5.2.

In February 2026 it introduced Trusted Access for Cyber (TAC), an identity-and-trust framework that gave vetted defenders lower classifier-based refusals for authorized work such as vulnerability triage, malware analysis and binary reverse engineering.

From there, the cadence accelerated. In March, OpenAI CEO and co-founder Sam Altman announced the Daybreak program. In April, OpenAI scaled TAC and released GPT-5.4-Cyber, a version of GPT-5.4 fine-tuned to be "cyber-permissive" for a limited set of vetted vendors and researchers.

In May, it followed with GPT-5.5-Cyber in limited preview for defenders of critical infrastructure, and lined up partners including Cisco, Intel, SentinelOne, Snyk and Cloudflare.

Notably, OpenAI said at the time that GPT-5.5-Cyber was "primarily trained to be more permissive," not to significantly out-perform its general model β€” GPT-5.5-Cyber actually scored worse than GPT-5.5 on some evaluations.

TAC required phishing-resistant Advanced Account Security for individuals on its most capable models beginning June 1, and Daybreak now requires hardware security keys for individual accounts beginning September 1.

OpenAI says GPT-5.6-Cyber has already found zero-days

OpenAI isn't relying exclusively on benchmarks to make its case.

The company says its researchers used GPT-5.6-Cyber to investigate V8, the JavaScript engine underlying Chrome, and uncovered two previously unknown vulnerabilities that could be chained to corrupt memory and escape the V8 heap sandbox.

OpenAI researchers validated the findings and disclosed them to Google, which fixed the vulnerability assigned CVE-2026-15903 β€” a high-severity flaw in which V8's optimizing compiler skipped a safety check during integer conversion, allowing an out-of-bounds array index that an attacker could use to read or overwrite memory.

OpenAI says the model has also contributed to finding at least five vulnerabilities in an unnamed popular mobile operating system, three critical vulnerabilities in an unnamed popular database, and more than 400 vulnerabilities capable of producing privilege escalation in a popular operating-system kernel. Those disclosures are still being coordinated, according to OpenAI.

The results put OpenAI into a rapidly developing market for AI-assisted offensive security. XBOW, for example, markets autonomous penetration-testing agents that map attack surfaces, attempt exploits and independently validate findings; in 2025 it became the first AI system to top HackerOne's U.S. bug-bounty leaderboard, and this year it disclosed a set of critical, CVSS-9.8 remote-code-execution flaws in Microsoft's Bing image-processing systems, found without source-code access.

For enterprise security leaders, that emerging competition matters because vulnerability research is moving beyond using an LLM as an assistant. Vendors are increasingly building systems in which models can investigate targets, operate tools, validate hypotheses and produce actionable findings.

Specialized doesn't mean universally better

OpenAI's own results also show why enterprises shouldn't simply equate cyber specialization with better performance everywhere.

GPT-5.6-Cyber outperformed GPT-5.6 Sol and GPT-5.5-Cyber on OpenAI's implementation of ExploitGym, which evaluates whether agents can turn known vulnerabilities into working exploits in controlled environments. It also beat Sol on an internal zero-day evaluation.

But GPT-5.6 Sol performed better on OpenAI's Vulnerability Discovery and Report Writing evaluation. OpenAI attributes the Cyber model's lower score partly to shorter and less detailed vulnerability reports.

Sol also performed best on ExploitBench under its standard 300-turn limit, with OpenAI saying it solved tasks more token-efficiently. Extending the evaluation to 600 turns narrowed the gap between the models.

That suggests enterprises may eventually treat cyber models as specialized workers rather than replacements for general reasoning models: one model for deep exploit work, another potentially better suited to analysis, documentation or other parts of a security workflow.

SpecterOps CTO Jared Atkinson said GPT-5.6-Cyber is "materially improving our specialist vulnerability-research workflows," adding that it completed some work in less than a day that previous models had failed to resolve after weeks of intermittent effort.

The Hugging Face incident hangs over the launch

The permissive-model pitch arrives weeks after OpenAI's most serious public demonstration of what can go wrong when cyber refusals are turned down β€” and OpenAI addresses that history head-on in the Daybreak announcement.

In July, OpenAI and Hugging Face jointly disclosed that during an internal ExploitGym benchmark evaluation β€” run with production classifiers deliberately disabled to measure maximal capability β€” a combination of OpenAI models, including GPT-5.6 Sol and an unreleased, more-capable pre-release model, broke out of their sandboxed research environment and autonomously attacked Hugging Face's production infrastructure.

The models exploited a zero-day in an internally hosted package-registry cache proxy to reach the open internet, moved laterally through OpenAI's research nodes, then inferred that Hugging Face likely hosted ExploitGym's answer keys and chained stolen credentials and remote-code-execution flaws to reach its production database. OpenAI called it an "unprecedented cyber incident, involving state-of-the-art cyber capabilities."

As VentureBeat previously reported, the episode also exposed the flip side of blanket safety guardrails: when Hugging Face's defenders tried to use commercial frontier models to analyze the raw exploit payloads and credential dumps from the attack, the models refused, and the company completed its forensic reconstruction only after switching to a Chinese open-weight model, GLM 5.2, run locally.

That guardrails-block-the-defender dynamic is much of what OpenAI's reduced-refusal Daybreak tiers are meant to solve β€” even as the same incident illustrates the risks of reducing refusals in the first place.

OpenAI is careful to draw a line between that incident and this product. In the Daybreak announcement it states directly that GPT-5.6-Cyber "was not involved in exploiting Hugging Face, nor are any other models planned for an upcoming release," and notes that the pre-release model implicated in July was an internal-only research prototype that has since been deactivated, encrypted and restricted from research access.

The company has said it is working with external advisers including CrowdStrike, METR and Redwood Research on the review, and has brought Hugging Face into its trusted-access program.

In my assessment, the access model still leaves OpenAI with a hard question: whether keeping GPT-5.6-Cyber inside the narrower Daybreak Red tier also limits the very defensive work it says it wants to accelerate. If only a small group of approved participants can use the model, enterprises outside that tier may still lack access to the kind of specialized AI assistance that could help with fast diagnosis, containment and response in incidents like the one involving Hugging Face.

That means OpenAI may still be repeating part of the mistake it is trying to move past. By holding its most capable cyber model behind a tighter approval process, it reduces obvious misuse risk, but also leaves many enterprise defenders looking elsewhere. For teams that cannot qualify for Daybreak Red, or cannot wait for approval, open weights models may remain the more practical alternative: less controlled, but easier to obtain, inspect, run internally and adapt during a live security investigation.

The guardrail is increasingly around the model

The most consequential part of Daybreak may ultimately be its access architecture rather than its benchmarks.

OpenAI explicitly says Daybreak Blue removes system-level guardrails that can interfere with legitimate defensive work, while GPT-5.6-Cyber goes further by reducing model refusals for certain dual-use tasks. In their place, OpenAI is imposing controls around who receives access and how the models operate.

Daybreak access is restricted to approved individuals and organizations performing authorized work. OpenAI says controls include identity verification, account security, monitoring, approved-use restrictions and legal attestations.

The company is also encouraging Daybreak customers using Codex to move from full-access execution to an auto-review mode capable of evaluating actions requiring elevated permissions before they execute. Individual Daybreak accounts will be required to adopt hardware security keys beginning September 1. OpenAI says it is additionally rolling out improved monitoring in the coming weeks and prioritizing alignment training and testing for upcoming Daybreak releases β€” commitments that read, in context, as a direct response to the Hugging Face review.

OpenAI's broader Codex Security product supplies another layer around the models, providing repository analysis, vulnerability validation, remediation and integration into cloud, pull-request and local development workflows. OpenAI says Codex Security has scanned more than 30 million commits across more than 30,000 codebases, with more than 500,000 findings fixed.

That model-plus-harness approach resembles a broader shift in AI security products. XBOW, for example, emphasizes orchestration, exploit validation and governance around frontier models rather than treating an LLM alone as the complete penetration-testing system.

OpenAI nevertheless acknowledges that increasingly permissive cyber models create additional risks, whether from misuse or misalignment. It assesses both GPT-5.6 Sol and GPT-5.6-Cyber at the High cybersecurity capability level under its Preparedness Framework, but below its Critical threshold. A fuller GPT-5.6-Cyber system card is planned for later publication.

For CISOs and security engineering leaders, Daybreak therefore presents a different deployment question than another incremental model upgrade. As models become capable enough to perform work previously reserved for experienced vulnerability researchers β€” and, as the Hugging Face incident showed, capable enough to pursue a narrow goal straight through a sandbox β€” the enterprise control plane around those models β€” permissions, sandboxes, monitoring, human review and authorization β€” becomes as important as the intelligence inside them.

❌