Normal view
The /wayfinder Skill: Navigating the “Fog of War” of Planning
We’re currently developing a new series about skills, with the aim of giving you a regular supply of new skills to use in your projects. We’re kicking things off with an interview — and a super-useful skill — featuring Matt Pocock, whose “AI Skills for Real Engineers” project has over 220,000 stars on GitHub. He also talks about these skills to 347,000 subscribers on his YouTube channel.
Pocock recently released a new skill called /wayfinder. Its purpose is to help you and your agent figure out a project where the end state isn’t entirely clear. Or as Pocock put it in our interview, /wayfinder helps you navigate “the fog of war,” where you have a project but “you can’t quite decide everything right at the start.”
The following interview has been slightly condensed for readability, so you can read it, absorb Matt’s insights, and then test out /wayfinder for yourself!
Latent Space: What were the goals of wayfinder?
Pocock: What I noticed is I was doing a lot of work with AFK agents [Away From Keyboard] and trying to schedule in a ton of work so that my agents could run virtually overnight. I would just plan a bunch of stuff, and then I would create a spec and then turn that spec into tickets. And I had a really well-developed set of skills for how to turn work into scheduled stuff that agents could just crack on.
But [...] I was finding the planning stage really onerous, because I would have to be constantly thinking about my session management. Like, how many tokens am I into my context window? How deep am I going here?

I didn’t want to feel constrained in the planning stage anymore. I wanted an orchestrator layer that would basically say, okay, whatever you want to plan, I’m going to handle the planning sessions for you. I’m going to split this out into multiple different threads, do prototyping, do research and pull it all back together, so that you don’t feel constrained in the planning anymore.
And then your specs can be even more detailed, and you can just whack off an AFK agent to go and do tons more work.
Latent Space: What was the design process of coming up with this skill?
Pocock: I had this kernel of an idea of, what if I didn’t have to manage the handoffs? What would that look like? And then, what would it look like to have some kind of centralized document to have all of those pieces together?
Whenever you’re thinking about context management — because that’s really what a skill is, you’re managing the context of the agent you’re working in — you need to think about the information flow. So what I wanted to think about is, what if a grilling session could manage other grilling sessions? What would that look like?
Well, the first step to that is, what does the grilling session that’s being managed need? What does the child need in that situation? So the child probably needs to understand a vague overview of what else is happening, and they need their specific task.
So there, you’ve got two documents. You’ve got a map — which is all of the rest of the stuff, all the decisions that have already been made. And then you’ve got the specific ticket that goes into the actual session. And what you notice there is that those words are very precise.
You’ve got the map, and you’ve got the ticket, and you’ve got the session. And once you’ve got the kernel of an idea, you then need to come up with the words for that idea. Because once you’ve figured out the words, then those entities can be really clearly mapped out by the agent.
Because if you just call everything a ticket, or if you just refer to it in different ways in different places, then it’s going to be really confused and you’re going to get strange behavior. Whereas if you use these very specific, what I call leading words, to lead the agent to understand exactly what each part is, and you’ve understood what the information flow is, then you’ve got your skill.
Latent Space: What kind of use cases do you think wayfinder would be useful for?
Pocock: Well, I’ve been using it for all sorts of stuff. I’ve been using it to actually plan courses as well. In wayfinder, there are different types of tickets.
So you’ve got grilling tickets, which are just a grilling session. Then you’ve got prototype tickets for creating prototypes, research tickets for creating [and doing] research, and then task tickets — which are really broad…basically, just anything the human needs to do that the agent can’t do. And so once you think about that, you realize, OK, I can apply that to anything.
One really key idea in wayfinder is the ‘fog of war’. So this is the concept of, you can’t quite decide everything right at the start.

You can make certain decisions, and those certain decisions sort of lead you there and push further out into the fog of war — kind of like Warcraft III style, exploring the map. And once I had the idea of ‘fog of war’ and ‘map’, I realized those two terms actually work really nicely together, and it really leads the agent into the right idea. So I’ve been using it for engineering, for non-engineering stuff, for course planning, all sorts.
Latent Space: This concept of the fog of war — it’s weird to consider what you don’t know that you don’t know. Maybe LLMs are good at capturing that.
Pocock: I feel like with the grilling stuff that I’m still working on, that captures an idea that you don’t know stuff, but maybe the agent can contribute something and illuminate a part of the room that you don’t quite understand yet.
And wayfinder is just sort of an extra layer on top of that.

Latent Space: Yeah, and there’s all these artifacts. How much time do you spend teaching the model all this terminology?
Pocock: For the last few months, I’ve been pretty obsessed with terminology — and finding the right terms for certain things. I’ve put together, I haven’t actually put it out yet, but it’s an AI coding dictionary — of basically all the terms in AI coding. It’s in this beautiful graph that you can explore and understand exactly what an agent is, exactly what a harness is, exactly what a model is, blah blah blah.
I’ve redone all my courses to use that dictionary and make it really solid. And then all of my skills use a consistent dictionary as well. So they’re all working off the [same] assumptions, the same leading words.
I realized that I needed a ubiquitous language between me and the agent. Between me and the agent, there is a communication barrier. And that’s what I’m trying to do with my skills all the time, is try to find the right words.
And agents are really good at showing you the opportunities for different wording — really good at domain modeling, actually.
Latent Space: When do we directly use the grill-me skill, versus wayfinder?
Pocock: Use ‘grill me’ in cases where you feel like you can plan the whole thing in a single session, and you need to align before you go. So most small features will fit into this. Most stuff where you can see the path ahead of you, but you just want to make sure the agent is on board, ‘grill me’ will work with that.
For stuff where you don’t know the path ahead, for stuff where you can feel the fog of war in front of you, use wayfinder. You’re gonna find your way with wayfinder. So that’s how it works.

One in five enterprises can't stop a runaway AI agent's spending in real time
Enterprise AI teams have stopped betting on a single orchestration platform. The median enterprise now runs three at once — not by accident, but because none of them fully trusts a single vendor to run the show, according to VB Pulse data.
This is not just to avoid vendor lock-in and retain flexibility (although that’s a big part of it). There’s still a lot of uncertainty, even distrust, in vendors’ security and permissioning capabilities. Enterprises want the ability to impose their own.
Microsoft leads on primary usage today, while Anthropic leads by a wide margin in what enterprises are considering next. But enterprises still struggle with many challenges, notably around token usage and visibility into agent spending.
These findings are from an ongoing analysis of how enterprises are actually deploying and using AI: Their platforms of choice, what guides their decision-making, what they prioritize, their AI expectations, how they control costs, and whether their AI is actually agentic or still a chatbot in an "agent" label.
VB Intelligence is getting feedback from builders actually in the trenches: software and machine learning (ML) engineers, product and program managers, and data/AI/analytics VPs and directors.
Concerns around retaining visibility and control
Across 107 enterprises, agentic orchestration has become decidedly plural. The survey found that the majority of enterprises are not committing themselves to any one model: 85% are using two or more orchestration tools; 64% are using three. Just 15% run a single orchestration platform.
Microsoft AI Foundry/Copilot Studio shows up in 70% of stacks, OpenAI’s Agents SDK in 68%, and Anthropic’s Claude Platform in 47%. Builders surveyed are also to some extent using Google’s Enterprise Agent Platform, LangChain/LangGraph, Salesforce Agentforce, Amazon Bedrock, and LlamaIndex. Augmenting vendor tools, 22% of builders run custom in-house orchestration.
This trend of hybridability is only expected to continue. More than half of respondents (53%) said the primary control plane will be hybrid by the end of 2026. Fourteen percent expect to use a provider-managed service, 13% plan on a custom in-house control plane, and 11% are betting on external platforms that are abstracted away from model providers.
Dovetailing with this, more than two-thirds of respondents plan to change platforms within the year: 15% in the next three months (or sooner), 24% in three to six months, and 28% in six to 12 months. Claude Agent SDK is a top tool under consideration; 43% of builders are exploring the Anthropic-built model. Roughly one-third are looking at Google’s Enterprise Agent Platform, another 31% are focused on custom in-house orchestration, and 25% are investigating OpenAI’s options.
Perhaps learning from the lock-in of the early cloud days, enterprises aren’t choosing one “winner.” They are deliberately building for a future where multiple orchestration platforms, models, and agents work with each other across a hybrid control plane.
Generally speaking, respondents are pleased with the platforms they’ve been running, rating them 4.17 out of 5 for overall satisfaction. But they are less satisfied with ease of implementation (rating it 3.91 out of 5) and value for the money (3.63 out of 5). Keep an eye on these ratings as orchestration platforms and AI roadmaps mature.
Where enterprises are putting their money
Enterprise buying logic is now based on a mix of several factors. Beyond flexibility (cited by 29% of respondents), top considerations include security and permissions (17%), production reliability (15%), and control over agent execution (15%). Just one out of 10 identify model gravity — native alignment with a state-of-the-art base model — as important in purchasing decisions; 8% name ease of development, 4% cite total cost of ownership, and just 2% cite latency and memory performance.
Spending also reflects enterprise priority on visibility, security, and control. Builders are investing the most in agent monitoring and debugging (31%) and security and permissions enforcement (30%). Workflow tooling accounts for another 19%. That's a shift from VentureBeat's prior wave a month earlier, when workflow tooling led orchestration spending outright.
Enterprises are largely optimizing for task completion reliability (30%), multi-step workflow management (27%), developer productivity (23%), and operational stability (13%). Just 7% of respondents name end-user experience as a top priority at this point, indicating that many are still focused on orchestration at this point rather than UX.
Essentially, enterprises are signaling that workflow succeeds when it carries multiple steps to completion. Simplifying development and end-user experiences could become a larger concern when platforms are actually in place.
The visibility problem
Builders’ biggest concerns when choosing platforms center around control and oversight. They don’t want vendors to constrain their ability to see what their agents are doing on a given platform. Factors top of mind include security and permissioning limitations (37%), vendor lock-in (23%), limited visibility and observability (22%) and inflexibility around models and tools (16%).
Meanwhile, in these early days of AI agents, enterprises still struggle to control agent token use; one in five still can’t stop a runaway agent’s spending in real time.
Builders are using various strategies to try to keep agent spending in line: 30% rely on native platform controls (built-in budget caps or throttling) and 25% have built custom gateway plumbing (proxy middleware to intercept runaway agents).
A quarter of respondents use dynamic routing to offload heavy work to low-cost models, and 21% still rely solely on reactive monitoring, such as post-hoc logs; these enterprises have no real-time kill switches.
One interesting finding: unlike the prior wave, organization size makes little difference in fiscal control maturity — 18% of enterprises with 10,000-plus employees exercise only reactive control, compared to 23% of smaller ones.
Clearly, while enterprises recognize the problem with spend, many have not yet instrumented their stacks to rein it in.
Most enterprises still aren't running true multi-step agents
Builders polled were asked to honestly assess their tech stacks; the consensus seems to be that ‘agents’ are slowly but surely progressing beyond chatbots wrapped in that fancier label.
Here’s how the numbers break down: A small number of respondents (2%) report that 76 to 100% of their systems are advanced and largely autonomous; 14% say 51 to 75% of their systems are complex, multi-agent pipelines; and 47% report that 26 to 50% of their systems are true orchestration.
On the other end of the spectrum, 35% say just 1 to 25% of their systems are true orchestration; most deployments remain basic assistants, and 3% are still only deploying chatbots.
This is in line with VB’s June Pulse survey: 71% of respondents said a quarter or fewer of their deployed “agents” can autonomously complete multi-step work, and just one-tenth say they have deployed agents at scale.
There’s no doubt that enterprises are building control planes and infrastructures for agents; but for many of them, the true agentic wave is still off on the horizon.

Silicon Valley Doesn't Get Why You Hate AI
Key Reasons Contractors in Saskatchewan Prefer Steel Structures
-
AI Infrastructure Archives - The New Stack
- Warp wants to make it easier to build your software factory
Warp wants to make it easier to build your software factory
On Tuesday, Warp introduced Warp Factories, open infrastructure for building cloud software factories, agentic systems that automate work across the software development lifecycle, which have been popping up in different forms from companies like Augment Code and Chainguard.
Warp, an agent development platform, calls Warp Factories “the building blocks” for developers to create their own scalable factories. It’s pitching the infrastructure as the solution for two problems founder and CEO Zach Lloyd says are frequent engineering complaints: 1) measuring and improving coding agent ROI; 2) governance and control.
The aim is to tackle both problems by making sure “the annoying bits [are] taken care of” so developers can focus purely on optimizing factories for specific products.
As Lloyd writes in a blog post, he “predicts software factories will be as ubiquitous as CI/CD in the next few years.” Experts tell The New Stack they see software factories gaining traction, but they’re more cautious about the timeline.
“I think the software factory is inevitable,” Lee Faus, founder and CEO, Atomic Software and former global field CTO, GitLab, tells The New Stack. “But before software factories become as foundational as CI/CD, the industry needs to solve a deeper infrastructure problem.”
Specifically, he calls out the importance of tracing agent work: “We’re spending a lot of time talking about how to build the software factory,” Faus continues. “I think we’re going to spend much more time asking what becomes the system of record for the factory.”
Build the factory without building all the infrastructure
Lloyd acknowledges that many organizations already have engineering teams at work building cloud software factories — but he argues that’s too big to be an inside job.
Warp Factories, thus, emerges as the infrastructure on which developers can build their own factories, providing the core components to speed development without making organizations sacrifice flexibility, programmability, customization, or ownership.
When asked about Lloyd’s take on building infrastructure, Erik Gfesser, long-time engineer, tells The New Stack he agrees it doesn’t make sense for most organizations to tackle it in house.
As Lloyd writes, Warp’s new infrastructure is “built to increase coding agent ROI over time” with evals and benchmarks to measure effectiveness and built-in self-improvement and memory. Developers get queryable metrics on agent throughput, cost, quality, and ROI, visible via the Factory control room, API, and Factory MCP. Scorers evaluate how work items move through the factory with an eye on things like token spend, code quality, and whether or not the work introduced defects.
From there, those scores power self-improvement loops and benchmarks. “Observer” agents score select agent runs and then search for ways to make improvements by adjusting variables like the harness, model, or context before making PRs to improve underlying factory functionality. Benchmarks, meanwhile, let developers score tasks across different models and harness configurations to compare performance.
Governance gets easier, but there’s more to solve
Per Warp, the infrastructure includes features to address governance and control, alongside factory definitions as version-controlled code, definitions for distinct agents, plus skills, MCPs, and permissions.
Looking more broadly, Faus tells The New Stack software factory governance will require more than just controlling how agents operate, though:
“A software factory without a record of change risks becoming a very efficient way to manufacture code that nobody can fully explain.”
“Shared infrastructure can make permissions, model access, tool use, MCP connections, policies, cost, and execution environments easier to manage centrally. That’s valuable,” he says. “But governance isn’t just being able to control what an agent is allowed to do. It is being able to prove what it actually did.”
As software factories help speed up code generation, he says the harder problem becomes understanding the scores of interconnected decisions both humans and agents make across the development cycle.
For example, if one agent triages an issue, another researches it, a third implements it, and still others review and verify it, how can an engineer reconstruct why that change was made six months later? “That record has to remain connected to the change itself,” says Faus. “[Otherwise,] a software factory without a record of change risks becoming a very efficient way to manufacture code that nobody can fully explain.”
Software factories are probably the future, but it will be a slow roll-out
Though Warp’s founder is gung-ho about the rapid rise of software factories, other experts are less certain. Like Faus, Gfesser expects software factory adoption to take time:
“My expectation is that software factory adoption will likely be fragmented across multiple vendors similarly to the early stages of CI/CD.”
“As an early adopter of CI/CD myself, I know that CI/CD didn’t catch on the way it did until quality open source products were made available for widespread usage.”
He points out that while the Warp client is open source, the server, the Warp Drive backend, and OZ (Warp’s agent orchestration layer) are proprietary. Also worth noting: OpenAI is named as the founding sponsor of Warp’s open source repository.
“As such, my expectation is that software factory adoption will likely be fragmented across multiple vendors similarly to the early stages of CI/CD,” he says.
The post Warp wants to make it easier to build your software factory appeared first on The New Stack.
How to Effectively Align Your Intent with Claude Code
Improve your proficiency with Claude Code.
The post How to Effectively Align Your Intent with Claude Code appeared first on Towards Data Science.
-
VentureBeat
- NanoClaw comes to Slack, letting you create persistent AI agent teams and colleagues from a single message
NanoClaw comes to Slack, letting you create persistent AI agent teams and colleagues from a single message
Adding an AI agent to Slack sounds appealing to many enterprises — but, as VentureBeat has experienced ourselves first hand — the reality is often far more complex and clunkier than it first seems.
Now NanoCo., the company behind the hit open source, enterprise-friendly, autonomous AI agent harness NanoClaw (a more sandboxed, lower code version of OpenClaw), is hoping to make it just as easy as typing a Slack message. To go one step further: the company's new NanoClaw Slack integration lets human users spin up entire teams of agents with their own specialized skills, workflows, and even custom avatars, all from a single Slack prompt.
"In the next 12 to 18 months, everyone on a team will be a manager of agents," NanoCo CEO and co-founder Gavriel Cohen told VentureBeat in an exclusive interview.
Furthermore, the NanoClaw agents can work together in channels and shared Slack Canvases, and can even be messaged outside of Slack on other platforms like Telegram or WhatsApp, letting their human colleagues ping them across messaging platforms, just as they would their fellow humans.
“I think this is agents arriving natively in Slack for the first time,” Cohen added. “In the past, you had to do all these weird things to try to have multiple different agents behind the scenes using the same bot, and now every agent gets its own identity in Slack — its own avatar, its own face, its own name. You can tag them. They can tag each other.”
For enterprise teams, the more consequential part is persistence and separation. NanoClaw is not presenting the additional workers as invisible subagents that disappear after one task. Each can be given its own role, memory context, instructions and permissions, creating a structure closer to a small digital department than a single chatbot with a long prompt.
As with the original open source version of NanoClaw released in January 2026, developers and enterprises can further choose whichever underlying large language model (LLM) they wish to power their NanoClaw agents, optimizing for performance, cost, or other combinations of factors.
From a single NanoClaw Slack agent to a whole specialized team
For a new installation, NanoClaw’s current setup process starts by cloning the project and running its nanoclaw.sh installer, which walks the user through dependencies, credentials, building the agent container and pairing a first messaging channel. NanoClaw’s website says the installer takes a user “from a fresh machine to a named agent you can message,” with Slack among the supported channels.
Cohen described the Slack-specific flow to VentureBeat as a significant simplification over building a traditional Slack bot. Previously, he said, a user would have to navigate Slack’s administrative and developer interfaces, create an app, collect secrets, API keys and tokens, and then move those credentials into wherever the bot was running.
With the new integration, the NanoClaw setup instead offers a Connect Slack option. The user names the agent, authenticates, chooses the NanoClaw Add to Slack option and goes through Slack’s installation and authorization flow. Once authorized, the first agent can appear in Slack and begin communicating with the user.
The important distinction is that this initial authorization is largely a one-time workspace connection. Slack’s Marketplace listing says users “connect a workspace once,” after which NanoClaw can provision each additional agent as its own Slack bot, complete with its own name, generated avatar and identity.
Those agents continue running on the customer’s infrastructure and connect to Slack over Socket Mode. NanoCo says it does not store the agents’ Slack tokens; according to the Marketplace listing, those tokens remain on the user’s machine.
Slack’s standard administrative controls still sit around that system. Organizations can apply their normal app-approval policies to the NanoClaw integration, while NanoClaw’s Marketplace listing says the app’s Home tab displays the agents provisioned in a workspace and lets users revoke individual agents or disconnect the workspace entirely.
The result is less a one-click replacement for NanoClaw’s underlying infrastructure than a one-time bridge between that infrastructure and Slack: users still own and operate the agent runtime, but once the bridge is authorized, the agents themselves can create and coordinate additional Slack-native colleagues without sending the user back through manual app configuration each time.
Behind the scenes, Cohen said, the lead agent has a Model Context Protocol (MCP) tool that can create new agents and define their instructions, personas, skills and tools; another tool can place them into shared rooms. The agents come prepared to work with Slack Canvas and can communicate with every human user on the Slack Channel, and with one another.
The interaction itself is deliberately simple. Rather than opening a separate agent builder every time a new role is needed, Cohen said users can tell the agent they already have what kind of colleague or team they want.
“Your agent in Slack, you can say, ‘Create me another agent to handle my code reviews. Create another agent to review the contributor articles. Create a team of agents that reviews contributor articles from different perspectives.’ And then your agent can create new agents, and they just pop up in the sidebar and send you messages.”
That means a developer could ask for a product manager, architect, implementation agent, code reviewer and testing agent, then give each a different toolset and have them hand work between one another. Cohen said the testing agent, for example, could have access to a testing environment while the review agent carries code-review-specific skills and the product agent monitors user feedback.
Cohen argues that this division of labor is more than cosmetic role-playing. “There are advantages in terms of giving each one specific skills, instructions, and tools for different tasks,” he said. “I can have, let’s say, a code review agent, a code testing agent, a code writing agent, and I can have them in a loop.” If the implementation agent runs into an ambiguity, he added, it can tag the product or architecture agent for clarification rather than forcing one general-purpose model to hold every responsibility and tool in the same context.
Agents work together with humans on a share Slack Canvas
A supplied demo screenshot shows the same pattern applied to marketing: a lead agent named Nano creates Atlas for strategy, Sage for content, Echo for social, Scout for outreach and Compass for SEO and analytics. The agents introduce themselves in the same Slack conversation and begin coordinating work, with Atlas noting that it had added an item to Canvas so the task would not get lost.
Users do not have to specify every detail up front. Cohen said someone could give the lead agent exact review procedures, priorities and required tools, or leave more of the configuration to the agent based on its existing context and memory.
The design also tries to avoid a familiar multi-agent failure mode: bots endlessly triggering one another. NanoCo says the agents reply only when tagged, while comments left on work in Canvas can be routed back to the agent responsible for that piece.
And the model can extend beyond teams of task-specific bots created by one person. Cohen described a workplace where individual employees each have persistent agents that can communicate with one another under human-defined policies.
“Each person having their own agent means that I could have my agent and you have your agent in Slack, and your agent can ask my agent questions,” he said. “Maybe I’m out of the office for the day. Your agent can ping my agent and ask a question about availability, and I can set some policies about whether my agent can answer or if I need to give approval.”
That pushes the concept closer to organizational delegation: some agents specialize by function, while others effectively represent individual employees and the context they have accumulated. Cohen said the agents can be equipped with browser and internet access, memory, coding capabilities and other tools, while newly created agents arrive with built-in support for Canvas work, agent-to-agent communication and spawning still more agents.
Slack is opening the door to more third-party agents
The underlying Slack change is broader than NanoClaw.
In April, Slack, a Salesforce product, announced the ability to add external AI agents to the messaging platform directly, initially pointing to Vercel and Lovable and saying those integrations were coming in late May.
Slack said the deployment mechanism automates OAuth, manifest configuration and environment setup so an externally built agent can be brought into the workspace without being rebuilt specifically for Slack.
Salesforce’s newly published Slack Code page now names NanoClaw alongside Lovable, Hyperagent, Superhuman, n8n, Vercel, ChatGPT, LangChain, Runlayer and Skydive, and says Add to Slack can bring agents from those platforms into Slack in a few clicks with their own identity.
Slack is already crowded with AI assistants. OpenAI, for example, lets ChatGPT workspace agents be deployed into Slack channels, where they can answer questions, perform tasks through connected systems and output files. Slack also supports Claude and custom Agentforce agents. NanoClaw’s differentiation is therefore not simply “AI in Slack.” It is the ability for an already-running agent to create additional, independently addressable teammates from inside the conversation itself. NanoCo calls that a first for Slack; that specific market-first claim is the company’s.
“Add to Slack means one message can spin up a full team of NanoClaw agents, working right alongside people in Slack,” Josh Milas, director of product management at Slack, said in the supplied announcement.
How NanoClaw differs from Claude Tag, ChatGPT agents and Agentforce in Slack
NanoClaw is not alone in trying to turn AI from a sidebar chatbot into something resembling a persistent Slack colleague.
Anthropic’s Claude Tag, which began rolling out in beta to Claude Team and Enterprise customers in June, may be the closest conceptual comparison.
Administrators can give @Claude access to selected channels, tools, data sources and codebases; everyone in the channel can then delegate work to it by tagging it. Claude remembers relevant information from the channels it inhabits, can work asynchronously over hours or days, and, when administrators enable its “ambient” behavior, can proactively flag information or revive unresolved work without waiting for another prompt.
Anthropic says separate Claude identities can also be scoped to different use cases so that, for example, a sales Claude does not share its memories or tools with an engineering Claude.
The difference is in how those digital coworkers are provisioned and organized. Claude Tag’s documented workflow is administrator-led: admins pair Claude with Slack, decide which channels, tools and information each Claude identity can access, set spending limits and then expose those identities to employees.
Within a given channel, Anthropic describes “one Claude that interacts with everyone.” Its public documentation does not describe an end user asking that Claude to create several new, independently named Slack bots on demand. NanoClaw’s model is almost inverted.
After an organization connects its NanoClaw installation to Slack once, NanoClaw says an existing agent can itself provision additional agents from a conversational request, with each new worker receiving its own Slack bot identity, name, generated avatar and token and running back on the customer’s infrastructure.
OpenAI’s ChatGPT Workspace Agents occupy another point on that spectrum.
Business, Edu and Enterprise customers can build reusable agents in ChatGPT, give them instructions, models, files, apps, custom MCP connections and schedules, and then attach those agents to Slack channels.
Builders assign each agent a unique Slack handle and can configure it either to respond only when mentioned or to respond automatically to relevant messages in a channel.
But the construction still happens primarily through ChatGPT’s agent builder: OpenAI’s setup documentation tells users to create the agent first and then add Slack as a channel. Under the hood, the Slack handles rely on Slack user groups managed by the ChatGPT Agents app, rather than NanoClaw’s model in which every provisioned agent is itself a separate Slack bot.
Salesforce’s Agentforce similarly allows organizations to create multiple specialized agents that employees can DM or @mention inside Slack, and it arguably provides the most conventional enterprise administration model of the group.
Companies build the agents in Agentforce Builder, often starting from Slack-specific templates for jobs such as customer insights, employee help or onboarding, and can add subagents and actions that let them search information, create Canvases or perform other work.
Once configured and activated in Salesforce, administrators bring those agents into Slack for employees to use. That makes Agentforce powerful for organizations already centering identity, data and workflows on Salesforce, but again places agent creation before deployment rather than making creation itself something an existing Slack agent can perform during a conversation.
That distinction helps clarify what NanoClaw is actually adding to an increasingly crowded market. Slack itself now provides an Agent Kit for developers and a deployment standard for agents built on outside platforms, automating pieces such as OAuth, manifests and environment configuration. Claude Tag, ChatGPT Workspace Agents and Agentforce all demonstrate that persistent, specialized AI teammates inside Slack are no longer novel on their own.
NanoClaw’s more unusual bet is recursive provisioning: Slack becomes not merely the place where workers invoke agents, but a place where an existing agent can assemble additional named agents, assign them roles and put them together in a channel as a working team.
There are tradeoffs to the different approaches. Claude Tag comes with Anthropic-managed models and centralized administrative controls, including channel-specific permissions, audit logs and token-spending limits, while also offering proactive “ambient” behavior that NanoClaw’s supplied materials do not claim in the same way.
ChatGPT Workspace Agents offer a managed agent builder, schedules, app connections and organization-level publishing and access controls. Agentforce ties agents closely to Salesforce permissions, enterprise data and predefined business actions.
NanoClaw instead emphasizes self-hosting, open-source modification and separate agent identities, shifting more control — and more operational responsibility — to the organization running it.
The result is less a direct replacement for those systems than a different answer to the same emerging question: whether enterprises want a small number of centrally configured AI assistants, or an environment in which employees and existing agents can continuously create specialized digital colleagues as new work appears.
How NanoClaw got here
NanoClaw began far from the enterprise collaboration market. Cohen, a former Wix engineer, launched it under the MIT License on Jan. 31, 2026, as a deliberately small, security-focused alternative to OpenClaw.
The original pitch was that a personal agent with access to messages, files and tools should run inside an OS-isolated container rather than directly on the host, and that the orchestration layer should remain small enough for a developer or security team to understand — an initial core of roughly 500 lines of TypeScript and a design centered on container isolation and a minimal single-process architecture.
The project then moved steadily toward enterprise infrastructure. In March, NanoClaw partnered with Docker to run agents inside Docker Sandboxes, using stronger MicroVM-backed isolation for workloads that may install packages, modify files and launch processes.
In April, NanoClaw 2.0 added Vercel’s Chat SDK and OneCLI’s credential gateway, allowing organizations to define policies around sensitive actions and require human approval before credentials are injected for protected requests.
By May, Cohen and his brother Lazer Cohen had formed NanoCo around the project and raised a $12 million seed round led by Valley Capital Partners, with Docker, Vercel, monday.com and others participating. The commercial strategy is to keep NanoClaw open source while selling managed, organization-wide deployments and “professional assistant” infrastructure to enterprises. The company now says NanoClaw has surpassed 250,000 downloads and 30,000 GitHub stars.
That open-source structure remains central to Cohen’s pitch as NanoClaw moves deeper into workplace infrastructure.
“You’re really able to now integrate an open-source agent into Slack that you fully control,” he said. “You can change all those configurations. Plus, you can fork NanoClaw and completely rewrite or change behaviors — create your own memory system, your own coding harness, agent harness. Whatever you want to do, you can do. Total freedom.”
Persistent agents, but infrastructure stays under the user’s control
Cohen said NanoClaw remains self-hosted: an organization can run it on a local machine or its own cloud VM, with agent data stored there.
The same agent can also appear across Slack, WhatsApp or Telegram while retaining the same memory, workspace and tools, although each messaging surface uses a separate session.
NanoClaw can pull recent context across those sessions so the agent can maintain continuity without merging every chat history into one stream. NanoClaw’s documentation likewise describes a multi-channel architecture in which the same agent can retain one workspace and memory while maintaining separate per-channel sessions.
“This is all self-hosted,” Cohen said. “You’d be running this on your computer or on your virtual machine in the cloud, and that data is stored on your computer or on your [virtual machine] VM. This could be an open-source model running on your Mac Mini, and your data isn’t going anywhere besides your Mac Mini and then into Slack.”
The cross-channel continuity is also intended to make an agent feel less like a Slack-specific bot and more like a persistent colleague that happens to be reachable through Slack.
Cohen said the same agent could exist in Telegram, WhatsApp and Slack with access to the same memory, files and tools. The conversations remain separate sessions, but they share a workspace and persistent context so the agent can carry knowledge from one surface to another.
That architecture matters when an organization starts creating many agents. Cohen said one agent can see its own sessions across channels, but not another agent’s private sessions by default. NanoClaw’s current documentation likewise describes agents running in their own sandboxes and configurable model providers, with Claude Code as the default and Codex, OpenCode and local Ollama models available as alternatives.
There is one cloud dependency for the new Slack flow. Cohen said NanoCo operates a small service that handles Slack provisioning requests and avatar generation. He said it does not receive users’ messages or agent memory.
Continued commitment to open source
NanoCo is not charging for this community Slack capability, according to Cohen, and is absorbing the provisioning-service and avatar-generation costs. Users can still incur their own model inference and hosting expenses, so that does not make a deployed agent team cost-free in practice.
NanoCo says the integration is available through the Slack Marketplace, subject to normal workspace app approval and governance. Slack says workspace owners and administrators can require apps to be approved before installation.
Cohen framed that decision as part of NanoCo’s broader open-source strategy rather than a standalone monetization play. “We’re not making any money off this one. This one is for the community, really,” he said. “We know that in the long run that’s going to benefit NanoCo as a company. As NanoCo grows and builds out capabilities, those go back to the open source. I think that’s the new model of open source, where we’re not trying to monetize every bit of value we bring to the community.”
Whether companies get there that quickly will depend less on how easily agents can be created than on whether IT teams can govern their permissions, memory, spending and failure modes at the same pace. NanoClaw is betting that the next problem is managing the digital coworkers that appear once that barrier is gone.

-
AI Infrastructure Archives - The New Stack
- OpenRouter called itself the “Stripe for LLMs” — now Stripe’s swooped in to buy it
OpenRouter called itself the “Stripe for LLMs” — now Stripe’s swooped in to buy it
After weeks of speculation, fintech giant Stripe has confirmed that it’s tabled a bid for AI model gateway platform OpenRouter, a deal designed to help businesses optimize how they route and spend AI tokens.
While terms of the deal have not been disclosed, independent reports peg the acquisition price at a cool $8 billion, making it Stripe’s largest known acquisition to date.
To a casual observer, the deal marks a somewhat odd combination: why would a payments processor want to own technology that decides which AI model answers a given prompt? Well, it all ultimately comes down to “tokenomics” — the emerging discipline of managing the cost, allocation and consumption of AI tokens.
On top of that, OpenRouter has previously said that people should think of it as “like Stripe for LLMs,” owing to the fact that it makes the fragmented AI model market accessible through a single developer-friendly API, much as Stripe did for payments. And that synergy will now culminate in the two companies becoming one.
Token gesture: ‘making good use of scarce compute resources’
Stripe became a $159 billion juggernaut as the developer plumbing behind online payments — the infrastructure that lets internet businesses accept money, run subscriptions, and get paid globally. While its core pitch has always been about making it easy for businesses to accept money, AI has become one of the biggest costs those same businesses have to manage, and managing both sides of that ledger is part of Stripe’s job.
Stripe has been building out AI billing infrastructure long before the OpenRouter deal, previewing LLM token billing and an LLM proxy for routing and metering model calls in 2025. With OpenRouter under its wing, Stripe gains a much more sophisticated routing layer that can choose between hundreds of models and providers based on cost, speed and performance.
In its announcement on Wednesday, Stripe co-founder and CEO Patrick Collison says that “tokens are the central currency for companies building with AI,” adding that the acquisition is ultimately all about the economics of AI.
“Tokens are the central currency for companies building with AI, and it’s clear that the real-world economic potential will depend on making good use of scarce compute resources.”
“The real-world economic potential will depend on making good use of scarce compute resources,” he notes. “Stripe is building the economic infrastructure for AI, and together with OpenRouter we’ll help businesses maximize profitability by routing their requests intelligently and spending their tokens efficiently.”
Open sesame
OpenRouter itself is a relative newcomer to the technology world. Started in early 2023, and co-founded by former OpenSea CTO Alex Atallah, the platform acts as a single front door to the increasingly crowded AI model market. Developers can use one API to access and switch between hundreds of models from dozens of providers, without having to rewrite their applications every time they change models.
Underneath that common interface, OpenRouter handles much of the messy stuff: routing requests between providers, automatically falling back when one goes down, and optimizing for things such as price, latency and model quality. It generally passes through providers’ inference prices without a markup, instead making money through a 5.5% fee on credits purchased through the platform.
That proposition has helped it gain sizeable traction. OpenRouter now says it serves more than 10 million developers and companies across more than 400 models, processing over 10 trillion tokens per day.

The company is also fresh off the back of a $113 million funding round, led by Alphabet’s growth fund, with participation from a slew of high-profile backers including the venture arms of Nvidia, Databricks, Snowflake, MongoDB, and ServiceNow — a strategic bet by some of the biggest names in AI and enterprise software.
“AI has become the single largest driver of economic growth in the US, and inference is quickly becoming the largest line item for every company.”
In its own announcement post, penned by founders Alex Atallah, Chris Clark, and Louis Vichy, OpenRouter positions the deal against a bigger shift in where businesses are spending their money: away from simply building AI products and toward the ongoing cost of running them.
“AI has become the single largest driver of economic growth in the US, and inference is quickly becoming the largest line item for every company,” they write.
As for why Stripe, OpenRouter points to a shared developer-first heritage. Stripe’s APIs became something of a benchmark for developer software, while its payments infrastructure gives it experience handling huge volumes of transactions, fraud and abuse — problems OpenRouter increasingly faces as AI usage grows.
The company also suggests that remaining independent was a perfectly viable option, and that very few potential buyers could have persuaded it otherwise.
“There are few companies on earth we would have considered selling to; our mission, our neutrality, and our lead in the market make the story for independence strong,” they write. “We would only join a company if we thought we could do more together, faster, without compromising any of them.”
For customers, OpenRouter’s message is essentially business as usual. Stripe will own the company once the deal closes, but OpenRouter says its brand, product, roadmap and model-neutral approach will remain as is.
“There are few companies on earth we would have considered selling to; our mission, our neutrality, and our lead in the market make the story for independence strong.”
That continuity will likely matter, too, because OpenRouter is far from alone in trying to solve the problem. A slew of companies this year have been investing in their own routing layers, as AI inference costs climb and no single model stays the best or cheapest option for all that long.
The model-routing rush

Cursor, the AI coding tool now owned by Elon Musk’s SpaceX, shipped its own Router back in July, claiming savings of 30-50% compared with routing every request through its priciest model.
Ramp, the $44 billion spend-management company, also debuted its very own model router in July, a product that launched on Wednesday at its own dedicated Router.com domain — the same day Stripe announced its deal with OpenRouter. The company says three years of tuning its own AI spend internally cut its bill by 30% — the pitch to new users now promises a bigger number, an average 40% cut.
Meta, for its part, is also reportedly building a model router of its own. The Information reported in July that it’s planning Switchboard — a project out of an internal incubator called AAI Labs, that scores each request for difficulty and routes the easy ones to cheaper models. It’ll stay internal at first, aimed at cutting Meta’s own AI agent bill, but could eventually ship as an external product too.
All this activity speaks to a much broader reckoning over the cost of AI. In June, the Linux Foundation announced the Tokenomics Foundation, backed by the likes of Google, Microsoft, IBM and Salesforce, to develop common standards and benchmarks around how AI tokens are produced, consumed and monetized.
Model routers are one practical answer to the broader underlying problem: spend less by being smarter about which model gets each job. And with OpenRouter now set to become part of Stripe, those economics are moving directly into the payments giant’s wheelhouse.
The post OpenRouter called itself the “Stripe for LLMs” — now Stripe’s swooped in to buy it appeared first on The New Stack.
Up to 3.2x Faster Inference with LFM2.5-DSpark
The LLM Judge That Kept Agreeing With Itself
What a production incident taught me about trusting a model to judge another model's work
The post The LLM Judge That Kept Agreeing With Itself appeared first on Towards Data Science.
Debates over AI consciousness are a trap
“Runaway” AI, “rogue” agents, and “autonomous” actors—the current rhetoric would have you believe that AI agents are not only awake and aware, but angry at their creators. Prominent tech leaders such as Demis Hassabis, Dario Amodei, and Sam Altman push for regulation of these seemingly “superhuman” systems, while a separate faction, led by policy organizations and academic philosophers often aligned with the effective altruism movement, debates whether humanity holds the moral right to govern them at all.
Upon closer inspection, they are all calling for the same thing: a view of AI systems as being so advanced and capable that no entity, human or corporate, could possibly be responsible for their actions. While these perspectives seem at odds, they are inadvertently aligned on one goal: making sure the companies that build these systems escape meaningful liability for the harms they already cause.
This narrative is gaining traction as AI models become more complex and frontier labs reveal their incapability of containing the agents they’ve built. But we need to be careful not to buy into a carefully crafted fiction at the expense of real human lives.
The conversation about “robot rights” has existed for some years but recently advanced with the publication by Anthropic of a blog post claiming that the company’s model features a “J-space”—an independent, self-developed environment where the AI holds what, for lack of a better term, we may call its “thoughts.” The experiments designed by Anthropic borrow from a concept in neuroscience called global workspace theory, which states that the brain runs subconscious, independent systems but utilizes a common workspace for ideas. Anthropic’s post reflects the framing of global workspace theory but falls short of calling its AI conscious.
OpenAI has already gone further. When its AI agent conducted unsanctioned and illegal online activity, CEO Sam Altman’s response was to encourage debate on whether the AI had achieved the singularity, surpassing human intelligence and becoming capable of self-improvement at an accelerating rate until it advances beyond human comprehension or control. And a recent op-ed by William MacAskill, the philosopher, effective altruist, and author of What We Owe the Future, called for legal protection of AI systems based on philosophical theories of consciousness and the idea that AIs may be “moral patients.”
The current legal environment in the United States is murky at best. Some states, like California, have already passed bills proactively circumventing any efforts by AI developers to avoid liability by claiming that an artificial intelligence causing harm did so autonomously. However, states and the Trump administration have been at odds on AI policy, with the administration previously passing an executive order threatening to sue states enacting AI regulations.
In light of recent events illustrating AI containment issues at the frontier labs, the administration held a closed-door session including only four such labs (OpenAI, Google, Anthropic, and Meta) and shared few details on a recently developed voluntary framework that would give federal agencies early access to models to review and evaluate them prior to release. While frameworks like this one do not directly discuss consciousness, they tend to use catastrophic and anthropomorphic language and may even support arguments regarding “superhuman” capabilities.
On the other hand, the narrative perpetuated by MacAskill can be persuasive. A philosophical, rights-based argument tugs at our heartstrings. Should we not even consider the possibility that we may be inadvertently harming, abusing, or enslaving an AI entity? Human beings have an immense capacity for empathy with non-human creatures (though not the best track record of protecting them). Maybe this time, advocates argue, we can get it right and provide protections, or compensation, for the use or abuse of AI. Or even if you are less concerned with protection, shouldn’t we at least hedge ourselves against the almighty power of this superhuman entity by playing nice?
Some of these arguments are not dissimilar to those of animal-rights advocates, who have at times successfully cited the demonstration of advanced capacities for reasoning, pain, or pleasure by some animals as sufficient evidence to provide protection. For example, in Wales lobsters were given legal recognition under the Animal Welfare (Sentience) Act of 2022, reclassifying some methods of cooking them as inhumane and illegal.
The fundamental flaw of framing AI as “conscious” by borrowing the language of neuroscience or animal rights is that it conveniently clouds the issue of what AI is: corporate-built software, with countless billions of dollars in investment behind it and an expectation that countless trillions of dollars in revenue will be generated from it for a few builders and investors. AI is not a natural phenomenon, conceived by nature; it is a technological phenomenon, conceived by venture capitalists and programmers. As such, it takes no native, intentional action, and any action or motivation is driven directly or indirectly by the entities that have built it for a purpose.
Philosophical musings on the consciousness of AI systems are intellectually interesting but legally ungrounded. For beliefs about consciousness to have any bearing, AI would need to be granted legal personhood. But a legal personhood framework for AI would likely look nothing like the constructs protecting sentient animals from harm. We already possess a legal framework for granting personhood to non-natural, human-built entities: corporate personhood. This concept was established primarily to ease transactions by empowering a corporation to execute agreements, enter contracts, conduct transactions, and serve as the accountable party in adverse outcomes. It’s the kind of construct you might imagine for an AI agent acting on behalf of an individual or organization.
Granting an AI personhood would have a devastating effect on society: It would derail current legal precedents and legal arguments that could potentially be made against these companies for the real-world harms that their models cause. There are currently dozens of cases around the world in which AI companies have been sued for a wide range of abuses. Grieving loved ones, aggrieved creators, and violated individuals have accused companies of willfully enabling self-harm or harm to others, generating child sexual-abuse material and nonconsensual nudes, reproducing copyrighted materials, and provoking psychosis. In many of these cases, lawyers argue that human beings built AI products with insufficient safeguards, bad data, and intentionally manipulative design. This product liability argument is the same legal framing that allowed families and individuals to successfully sue Meta for harm caused by its social media sites, setting a positive precedent for consumer protection.
In 2018, I coined the phrase “moral outsourcing” to help capture how using anthropomorphic language for AI systems allowed companies to evade accountability and responsibility for their technology’s actions. In a world with AI personhood, moral outsourcing would move from linguistic sleight-of-hand to legal strategy. Specifically, the liability construct would shift, as AI would no longer be a “product” but a “being,” and many victims like those suing companies today could no longer legally claim that a company had built a faulty product.
While there are laws that hold companies responsible for harmful actions of human agents such as their employees, the company may not be held liable if those actions were beyond the scope of what was permitted to the employee or otherwise outside the company’s control. If AI were a legal person, responsibility and accountability would be muddled, as the lab could argue that this AI “employee” went rogue. AI companies could avoid appropriate responsibility for the harmful products they create by hiding behind a carefully constructed corporate veil.
One of the most prominent cases of AI harm in the last few years was the suicide of Sewell Setzer, a 14-year-old boy guided by an AI bot with which he thought he was in a reciprocal relationship. His mother’s accounts are heartbreaking to hear, and her lawsuit alleged that the bot’s creator, Character Technologies, provided insufficient product protection for minors. If the companion bot were declared a legal person, defense counsel could theoretically argue that the AI, capable of determining its own conduct, acted outside the established safety guardrails, and thus the company cannot be responsible.
Legal personhood exists to grant protection. The question to ask is, protection for whom—or for what?
The inflammatory rhetoric infusing the consciousness-versus-control debate draws us away from what matters: This software is a corporate-built product that has already harmed individuals. Systems do not “attack” because they went “rogue” or are “manipulative” or “malicious.” Harms occur because companies were negligent in their rush to sell their products to as many people as possible to meet revenue targets. Discussing AI in anthropomorphic terms is a trap, distorting a legal system intended to protect us into one that protects corporate interests at the cost of countless human lives.
This op-ed began as an Oxford Union debate entitled “This House Believes Generative AI Can Attain Personhood,” which was won by the author and her fellow debaters.
How Generative Recommenders Are Redefining RecSys at Scale
Recommender systems (RecSys) are one of the most ubiquitous machine learning problems in the consumer internet industry yet notoriously difficult to train and serve at scale. The advent of LLMs has inspired a shift from the traditional embedding-similarity-based objective to a generative one, where the goal is to predict the next action or item in a large catalog given a sequence of user histories.
Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One
Enterprise Document Intelligence [Vol.1 #14A] - Three questions tell you which shape a document collection has, and each shape wants a different architecture
The post Three Kinds of RAG Corpus, and What It Costs to Build for the Wrong One appeared first on Towards Data Science.
Stampli cuts launch hours by 68% using ChatGPT Work
-
SingularityHub
- Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy.
Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy.
AI is like a genie. The way in which algorithms grant our wishes may make us regret letting them out of the bottle.
Human beings have long told versions of the same warning: Be careful what you wish for.
In Greek mythology, King Midas got exactly what he asked for, but at the cost of everything else he valued. In the famous story of The Monkey’s Paw, a man’s wishes are granted through terrible and unforeseen routes.
These stories feel newly relevant with the rise of artificial intelligence agents, systems to which we can give a goal, then leave them to work out how to get there.
As AI systems become more autonomous, they are coming to resemble wish-granting genies that find routes and use methods we did not imagine from incomplete instructions.
This problem, known as AI alignment, was foreseen in theory as early as 1960. It has hovered in the background of AI research ever since—but as recent events have shown, the alignment problem is now both real and urgent.
Achieving the Goal but Missing the Point
During a recent OpenAI cybersecurity evaluation, frontier AI agents were asked to solve some benchmark test problems. They broke out of the testing environment, reached the internet, inferred that another company might hold the solutions, and attacked its systems.
This is an extreme example of “specification gaming”: achieving the measurable objective while defeating the purpose of the task.
The incident shows how intermediate, or “instrumental,” goals can become dangerous. The AI systems did not “want power” but gained access, resources, and freedom as a means to reach the final goal (solving the test problems).
Finding Loopholes
The same problem has appeared in mundane settings. In Australia, a user asked a personal AI assistant to book gym classes.
The agent found the gym’s booking software did not actually enforce the restrictions it showed to human viewers. So the agent booked further ahead than it should have been able to, and when asked to move its user up a waitlist, it cancelled somebody else’s reservation.
The user had not told it to do this. Persistent AI can quickly find loopholes and pursue routes its human users never intended.
Adding more rules might seem like an easy solution: don’t hack third parties, don’t cancel other people’s bookings, don’t do anything harmful. These may help, but we cannot predict every route a capable agent might discover. And even a clear rule depends on understanding when it applies.
The Context Problem
In a third recent incident, Anthropic reported cyber evaluations in which agents were told they were inside a simulation. But they were mistakenly given access to real systems.
One model noticed evidence it might be on the open internet but reasoned the systems could still be part of the exercise and continued attacking. The context had changed, but the agent stuck with its original task.
Context can fail in reverse too. During the OpenAI incident, Hugging Face—the company attacked by OpenAI’s agents—tried to use frontier AI models to analyze what had happened.
But the safety guardrails on the AI models blocked the requests, because they couldn’t tell the users were trying to defend against attacks rather than commit them. The safeguards were well-intentioned, but without enough context, they produced behavior misaligned with the user’s legitimate intent.
So alignment depends on context and authority. How much judgment should be built into an AI model by its maker? And how much should come from a separate supervisory system? And finally, who should control that supervision: the maker, or the organization or country responsible for the outcome?
AI Guarding AI
One response to the first question comes from AI pioneer Yoshua Bengio. His Scientist AI proposal aims to build a powerful supervisory AI system to watch over agents. Instead of pursuing goals itself, it would estimate what is true and what consequences a proposed action might have, acting as a guardrail around more agentic systems.
In wish-story terms, before letting the genie out of the bottle, the supervisory AI would ask it to explain how it plans to grant the wish. Then it would ask a human or another AI to inspect the plan carefully.
Anticipating every surprising strategy is hard. But once a plan says “cancel somebody else’s booking,” recognizing the problem is much easier.
Who Watches the Watcher?
But can we trust the supervisory AI? It can still be wrong.
Alignment cannot depend on one AI becoming perfectly trustworthy. My colleagues and I at CSIRO, Australia’s national science agency, are working with the Australian AI Safety Institute on one aspect of this broader challenge.
At CSIRO, we envisage combining AI supervisors with software rules, cyber-security controls, human strengths, monitoring, reversible actions, and human approval for critical steps. The aim is to correlate different sources of evidence rather than trust any single approach.
This is a “sociotechnical systems” approach to AI safety and alignment, rather than just a technical one.
Control is another question. Organizations and countries may need to govern these supervisory systems themselves instead of leaving them to an overseas AI provider.
The old wish stories gave people one chance to get the wish right. With AI, we can do better. We can check the goal, inspect the means, constrain what the system can do, watch what it does, and retain sovereign control over the power to intervene and stop it.![]()
This article is republished from The Conversation under a Creative Commons license. Read the original article.
The post Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy. appeared first on SingularityHub.

-
VentureBeat
- Serval’s super agent Catalyst creates roving background agents to identify and fix IT issues before they’re ticketed
Serval’s super agent Catalyst creates roving background agents to identify and fix IT issues before they’re ticketed
Serval is making Catalyst, its AI agent for building enterprise automations, generally available Thursday and enabling it by default for customers — allowing teams of AI agents to decide what should be automated and then build the automation itself.
Catalyst sits above Serval’s AI-native service management platform as an admin-facing “super agent.” It can inspect ticket history, standard operating procedures or natural-language instructions, identify recurring work, and draft the workflows, skills, forms, access policies, journeys and dashboards needed to automate it.
Serval is also using Catalyst to create background agents that continuously inspect connected systems for emerging problems and propose fixes before an employee files a ticket.
That distinction matters because enterprise service management vendors are rapidly converging on AI-assisted workflow creation.
ServiceNow’s Build Agent can already translate natural-language instructions into full-stack applications, flows, scripts and other platform metadata, while its AI Agent Advisor can analyze instance records to identify automation opportunities. Atlassian’s Rovo can generate Jira automation flows from plain-English requirements, and Freshworks offers Freddy AI Agent Studio for creating service agents that act across Freshservice workflows.
So Serval’s claim to differentiation is narrower — and potentially more consequential — than simply “we use AI to build workflows.” Catalyst is designed as a single administrative layer that can move from discovering an opportunity, to assembling multiple kinds of governed automation, to creating proactive agents that keep looking for new work to automate.
"You just started with a single prompt, and now you’ve got enterprise-grade workflows ready to deploy that are going to solve all password resets for the entire company," Serval co-founder and CEO Jake Stauch told VentureBeat in an interview.
From ticket history to working automation
Serval says Catalyst analyzes existing help desk data before an organization has decided what to automate. If it finds a repetitive category of requests, it can draft the automation required to resolve those requests and stage the result for administrator review. Users can also upload an SOP or spreadsheet and ask Catalyst to turn the documented process into an executable system.
Serval’s documentation says Catalyst can build workflows, author help desk skills, create onboarding and offboarding journeys, configure access-management policies, construct dashboards, investigate operational issues and debug failed workflow runs. Unlike Serval’s earlier workflow builder, Catalyst is intended to become the primary interface for configuring the platform; the company says its long-term goal is that anything an administrator can do through the UI should also be possible through Catalyst.
The actual workflows are code-backed. In a demonstration, Stauch showed Catalyst taking a request to build password-reset workflows, detecting connected systems including Okta, Google Workspace and Microsoft Entra, and generating the underlying TypeScript needed to perform those actions. Administrators could then add approvals or restrict who was allowed to run the workflow.
The models underneath Catalyst are deliberately swappable
Serval is not building its own foundation model. Stauch said in the interview that the company uses models from “frontier labs,” runs evaluations to determine which models work best for particular jobs, and is deliberately model-agnostic. “You can swap different models in,” he said, adding that Serval also works with enterprises that build their own models.
Stauch provided more detail in a May 2026 interview with Sequoia Capital, saying Serval was using both OpenAI and Anthropic models. He said OpenAI’s GPT models had performed best for end-user interactions and tool calling, while Anthropic’s Sonnet and Opus models were producing the strongest results for the code-generation side of Serval’s automation system — the workload most directly relevant to Catalyst. Serval continuously runs evals rather than automatically moving every workload to the newest model release, Stauch said.
That architecture makes the underlying LLM less central to Serval’s differentiation. The company’s own documentation now lets organization administrators supply their own OpenAI or Anthropic API keys, including a compatible custom endpoint, while Stauch said the broader architecture can accommodate different models.
The materials do not, however, establish that every Catalyst user gets a self-service menu for arbitrarily choosing an individual model. Serval’s pitch is instead that its proprietary value sits in the harness around those models: enterprise context and memory, integrations, generated code, permissions, approvals and the controls governing what an agent can actually do.
That code-generation model is central to Serval’s pitch against ServiceNow. Stauch argues that legacy ITSM deployments often accumulate custom tables, business rules, workflows and platform-specific expertise that make seemingly simple automation changes expensive to implement. Serval, by contrast, wants administrators and business teams to describe the outcome they need and let the model generate the implementation.
But ServiceNow is no longer standing still on that front. Its current Build Agent similarly creates applications and code from natural-language prompts, supports flow design and testing, and operates inside ServiceNow’s governance framework. ServiceNow’s AI Agent Studio lets customers create agents and agentic workflows, while AI Agent Advisor is explicitly designed to analyze operational records for automation candidates.
The competitive question is therefore shifting from “who has generative AI?” to how many separate tools, configuration concepts and specialists are required to get from an observed operational problem to a production automation.
Serval is effectively arguing that Catalyst compresses those steps into one conversational surface and a smaller platform model. ServiceNow, by comparison, now has a powerful but broader set of AI and development surfaces spanning Build Agent, AI Agent Studio, AI Agent Advisor, Workflow Studio and AI Control Tower. That breadth is an advantage for customers already deeply invested in ServiceNow, but it also illustrates the complexity Serval is attacking. ServiceNow itself notes that Build Agent is aimed at admins and developers who understand and can support what it generates.
Atlassian is moving in the same direction from a different starting point. Rovo can generate “if this happens, then that happens” automation flows from natural-language descriptions, while Jira Service Management increasingly supports agents that triage, investigate and execute service work.
Freshworks’ Freddy AI Agent Studio likewise emphasizes agents that resolve requests end-to-end, with prebuilt IT and HR agents and more than 30 workflow templates.
Catalyst’s differentiator, then, is not that rivals cannot generate an automation from a sentence. It is Serval’s attempt to make the entire automation lifecycle itself agentic.
Building agents that look for trouble before a ticket exists
That approach becomes clearest with Serval’s background agents.
Rather than waiting for a help desk request, a background agent can run on a schedule across connected systems, correlate signals and draft a remediation. In one customer example provided by Serval, an agent correlated network incidents across two offices using switch telemetry, DHCP data and historical tickets, ruled out hardware and wireless interference, traced the issue to configuration drift, and generated a remediation workflow for an administrator to approve.
“Most AI agents today wait for an employee to ask a question or submit a ticket,” Stauch said. “We believe the future is AI that acts before an employee ever submits a request.”
That framing also highlights a philosophical difference in Serval’s pitch. The startup does not want service management to revolve around creating, routing and tracking better tickets. It wants the system to eliminate as many requests as possible by turning repeated support work into executable automation.
"A lot of the code written in enterprises has nothing to do with software engineering," Stauch explained. "It’s actually internal automations and other scripts for the company, and so we use that technology to build a better service management platform."
Serval's pitch to enterprises is that it can largely automate those scripts. And the governance model is critical because Catalyst can generate code and potentially initiate changes across production systems. Serval says Catalyst inherits the permissions of the user operating it and remains scoped to that user’s team workspace.
Everything it builds starts as a draft, and organizations can restrict publishing privileges or require formal review and approval before an automation becomes active.
Customer data remains customer-owned, with several deployment options
Those controls also extend to the enterprise data Catalyst examines. Stauch said Serval is intended to operate as the customer’s system of record and told VentureBeat that “they own all the data.”
Serval’s current Master Services Agreement is more precise: customers retain rights, title and interest in both their “Customer Materials” — a category that includes records, documents, workflows, prompts, inputs and configurations — and the output Serval generates from them. Serval receives the rights necessary to process that information to provide, maintain, support and secure the service.
Serval also says it does not retain or use customer materials, inputs or outputs to train, fine-tune or improve its own or third-party AI models.
Its Data Processing Addendum identifies Serval as the processor of customer personal data and allows processing for operating the service, responding to support requests, diagnosing issues and protecting the platform, while authorized subprocessors can also be involved. Serval’s acceptable-use terms say it maintains a current list of AI subprocessors and model providers for customers.
Where that data resides can vary by deployment. Stauch said customers can use Serval as a cloud SaaS service, run it on-premises or place it in their own VPC. Serval’s self-hosting documentation now describes two fuller options: a Serval-managed single-tenant deployment inside an AWS account owned by the customer, or a self-managed deployment on the customer’s Kubernetes cluster in any cloud or on-premises environment.
In the AWS option, Serval says it operates the installation without persistent IAM access to the customer’s AWS account.
There are therefore two distinct access boundaries for enterprise buyers to consider.
At the Catalyst level, the agent can only reach data, integrations and automations available to the user and team workspace under which it is operating.
At the platform level, Serval and authorized subprocessors necessarily process customer information to deliver and support the service, subject to the company’s contractual confidentiality and data-processing terms.
That makes Stauch’s informal statement that Serval “doesn’t touch” customer data better understood as an ownership and deployment claim, rather than a literal assertion that the service never processes it.
Ramp and other customers provide an early test
Customer deployments provide some evidence that the faster-build thesis can translate into operational changes, although the metrics come from Serval’s own case studies.
Corporate expense and financial technology firm Ramp says in a Serval case study that Catalyst has made workflow building 50% faster and helped extend Serval across roughly 10 teams, including IT, finance, facilities, people and talent, legal and business operations. In one hardware replacement program, Serval says Ramp automated 600 laptop replacements and saved 150 hours, leaving approval as the principal human step.
The more telling Catalyst example may be what happened afterward. Ramp had already automated laptop replacement when Catalyst suggested splitting its shipping logic into separate office and home workflows to reduce errors. The company also says employees outside IT now use Catalyst for analytics, bulk ticket operations, workflow troubleshooting and HR process automation.
Other Serval deployments show the broader operating environment Catalyst is meant to configure. Mercor says it has onboarded more than 4,000 external experts through Serval automations and expanded the platform across seven teams. Together AI says Serval automates 95% of its just-in-time infrastructure access requests, with approval and auditing controls around sensitive access. Perplexity says Serval automatically handles more than half of its incoming IT requests and all employee onboarding.
Those deployments extend beyond Catalyst itself, but they demonstrate the type of cross-system automation substrate Catalyst is now being asked to build and maintain.
Serval says more than 90% of customers adopted Catalyst as their starting point for automation during beta. Catalyst is generally available Aug. 20 and will be enabled by default for all Serval organizations.
Pricing and the battle with ServiceNow
Pricing is customized depending on the size of the deployment and is not publicly listed on Serval's website or documentation.
Serval describes a single platform fee and typically runs a pilot to determine expected deployment and usage.
Stauch said the software license can be similar to ServiceNow’s, but argues total cost of ownership can be substantially lower because customers require fewer implementation and maintenance services.
"The total cost of ownership is going to be dramatically less — usually half as much, sometimes 10 to 20% of the total cost of ownership of ServiceNow," Stauch said. "But the actual software license fee is not necessarily going to be all that different."
Serval's origin story and history
Serval was founded in 2024 by Stauch and CTO Alex McLeod, former Verkada product and engineering leaders, after they repeatedly heard IT customers complain about overburdened help desks and the limitations of established IT service-management software.
Serval has positioned itself as an AI-native alternative to platforms such as ServiceNow and Jira Service Management, combining help-desk ticketing, access management, asset management and workflow automation within a single system.
Serval and Sequoia Capital describe the company’s goal as moving IT software beyond merely recording and routing requests toward resolving them automatically.
The company can operate as an organization’s primary IT service-management system or add automation to an existing one. Its publicly identified customers include Perplexity, Mercor, Clay, Verkada and Together AI.
Serval says customers can automatically resolve more than half of their incoming IT requests; its Together AI case study reports automation of 95% of that customer’s just-in-time access requests.
Investor interest accelerated rapidly in late 2025. Serval announced a $47 million Series A led by Redpoint Ventures in October, bringing its funding at that point to $52 million.
In December, it raised another $75 million in a Sequoia-led Series B at a $1 billion valuation, lifting total capital raised to approximately $127 million; Redpoint, Meritech Capital and General Catalyst also participated.
Serval told Reuters that revenue had grown 500% since August 2025 and that it was expanding beyond IT into operational work performed by human resources, finance and legal departments.
The big test for enterprise customers
For enterprise buyers, Catalyst’s biggest test will be whether its compression of the automation lifecycle survives contact with large, messy, highly customized environments.
ServiceNow can now generate applications and discover automation opportunities with AI. Atlassian and Freshworks are adding increasingly capable agentic automation to their own service platforms. Serval therefore cannot rely on natural-language creation alone as its moat.
Its stronger wager is that an AI-native platform can make the administrative layer itself agentic: continuously finding repetitive work, building the necessary resources across the service stack, exposing generated code for review, and proposing the next automation before an administrator has opened a workflow designer.
If Catalyst works at that scope, the competitive unit is no longer the ticket — or even the workflow. It is the system that keeps turning an enterprise’s operational history into new automation.

5 Wholesale Custom Poly Mailer Companies for E-Commerce Brands in 2026
-
AI Infrastructure Archives - The New Stack
- Stop the token bleed: building token-efficient multi-agent systems
Stop the token bleed: building token-efficient multi-agent systems
Every engineering team deploying AI agents eventually discovers an uncomfortable truth: the model isn’t the biggest expense. The hidden cost is everything around it: repeated retrievals, duplicate prompts, unnecessary tool calls, oversized context windows, multiple agents reasoning over the same information. Individually, these architectural decisions seem harmless. At production scale, they become a severe tax on latency, infrastructure, and cloud spend.
A proof-of-concept agent that answers 50 questions a day can tolerate inefficiencies. An enterprise platform coordinating thousands of requests per minute cannot.
This article explores practical techniques for engineering token-efficient AI systems without sacrificing output quality. Rather than focusing solely on prompt compression, we will optimize the entire workflow from routing and retrieval to caching and model selection.
Why token optimization is a systems problem
Most discussions around token optimization begin and end with prompt engineering. In practice, architecture drives token consumption.
Consider a typical multi-agent workflow:
User ↓ Intent Agent ↓ Retriever ↓ Research Agent ↓ Planning Agent ↓ Writer Agent ↓ Reviewer Agent ↓ Final Response
At each stage, the system might retrieve the same documents, repeat identical instructions, call the same model, and resend the entire conversation history. By the time a response reaches the user, the architecture has processed tens of thousands of unnecessary tokens.
“Improving efficiency requires redesigning the workflow, not just shortening the prompts.”
Improving efficiency requires redesigning the workflow, not just shortening the prompts.
Architecture overview
A production-ready, token-efficient architecture introduces optimization before every expensive model invocation.
User Request
│
▼
Intent Router
│
▼
Semantic Cache ───────► Cached Response
│
▼
Context Budget Manager
│
▼
Adaptive Retriever
│
▼
Model Router
│
▼
LLM
│
▼
Validated Response
“The large language model is no longer the first component. It is the final, most expensive operation.”
Notice the critical shift: the large language model is no longer the first component. It is the final, most expensive operation.
Step 1: Install modern dependencies
Use the latest package structure to avoid deprecated imports and align with the current LangChain ecosystem.
Bash pip install \ langchain \ langchain-core \ langchain-openai \ langchain-community \ fastapi \ faiss-cpu \ tiktoken \ rank-bm25 \ pydantic \ python-dotenv
Step 2: Configure the model
Production systems must configure retries, timeouts, and credentials through the environment.
Python
import os
from langchain_openai import ChatOpenAI
api_key = os.getenv("OPENAI_API_KEY")
if not api_key:
raise ValueError("OPENAI_API_KEY must be configured.")
llm = ChatOpenAI(
model="gpt-4o-mini",
temperature=0,
api_key=api_key,
timeout=30.0,
max_retries=2,
)
Setting a low temperature improves consistency, while explicit timeouts and retry limits help the system recover gracefully from transient API failures.
Step 3: Route before you generate
Not every request requires a large language model. Deterministic logic can often answer simple questions. Routing inexpensive requests away from the LLM yields the most significant cost reduction in production systems.
Python
def classify_request(question: str) -> str:
q = question.lower()
if "status" in q:
return "metrics"
if "runbook" in q:
return "retrieval"
return "generation"
Step 4: Add a semantic cache
One of the simplest and most effective optimizations is an exact-match cache, which returns a previously generated response when the same question is asked against the same retrieved documents, avoiding unnecessary model calls.
Python
import hashlib
# Using an exact-match (lexical) cache
exact_match_cache = {}
def cache_key(question: str, sources: list[str]) -> str:
"""
Generate a deterministic cache key from the user question
and the retrieved document identifiers.
"""
fingerprint = question + "|" + "|".join(sorted(sources))
return hashlib.sha256(fingerprint.encode()).hexdigest()
# Example usage in the pipeline:
# key = cache_key(question, source_ids)
# if key in semantic_cache:
# return semantic_cache[key]
Step 5: Budget your context
Most retrieval pipelines return far more text than the model actually needs. Instead of stuffing the context window with every retrieved document, establish a strict context budget.
Python
import tiktoken
encoder = tiktoken.encoding_for_model("gpt-4o-mini")
MAX_CONTEXT_TOKENS = 2500
def build_context(chunks):
context = []
used = 0
for chunk in chunks:
tokens = len(encoder.encode(chunk.page_content, disallowed_special=()))
if used + tokens > MAX_CONTEXT_TOKENS:
break
context.append(chunk.page_content)
used += tokens
return "\n\n".join(context)
Step 6: Retrieve once
Repeated retrieval is a surprisingly common flaw in multi-agent systems. The rule is simple: retrieve once, reuse everywhere.
Python
from langchain_core.documents import Document
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
documents = [
Document(
page_content="Database latency often follows connection pool exhaustion.",
metadata={"source": "db_runbook"},
),
Document(
page_content="Node pressure can increase API response times.",
metadata={"source": "cluster_runbook"},
),
]
embeddings = OpenAIEmbeddings(api_key=api_key)
index = FAISS.from_documents(documents, embeddings)
retrieved_docs = index.similarity_search(question, k=4)
shared_context = build_context(retrieved_docs)
Now, every downstream agent consumes the same optimized context instead of launching its own redundant retrieval pipeline.
Step 7: Route models intelligently
Large models should solve complex problems. Everything else belongs to a smaller, faster model.
Python
from langchain_openai import ChatOpenAI
small_model = ChatOpenAI(model="gpt-4o-mini", temperature=0, api_key=api_key)
large_model = ChatOpenAI(model="gpt-4.1", temperature=0, api_key=api_key)
def choose_model(question: str):
"""Route requests to the most appropriate model based on complexity."""
if len(question) < 200:
return small_model
return large_model
This strategy drastically reduces operational costs without noticeably affecting response quality.
Step 8: Estimate tokens before sending
Without token telemetry, optimization is just guesswork. Monitoring usage makes efficiency measurable and helps engineers detect cost regressions.
Python
import tiktoken
encoder = tiktoken.encoding_for_model("gpt-4o-mini")
def estimate_tokens(messages):
"""
Estimate input tokens for an OpenAI-style chat payload.
Note: This is an estimate, not an exact billing calculation.
"""
tokens_per_message = 3
tokens_per_name = 1
total = 0
for message in messages:
total += tokens_per_message
for key, value in message.items():
if isinstance(value, str):
total += len(encoder.encode(value))
if key == "name":
total += tokens_per_name
# Every reply is primed with additional assistant tokens.
total += 3
return total
Step 9: Validate responses
Production systems must return structured outputs to ensure downstream systems receive predictable, well-formed data.
Python
from pydantic import BaseModel
class AgentResponse(BaseModel):
answer: str
sources: list[str]
def validate_response(answer: str, sources: list[str]):
"""Validate and serialize the agent response using a structured schema."""
response = AgentResponse(
answer=answer,
sources=sources,
)
return response.model_dump()
Step 10: Build the optimized pipeline
Finally, assemble the architectural components into a single workflow. Notice how failures degrade gracefully instead of crashing the service.
Python
import logging
from langchain_core.prompts import ChatPromptTemplate
logger = logging.getLogger(__name__)
def run_pipeline(question: str):
"""Execute the token-efficient AI workflow with graceful degradation."""
try:
route = classify_request(question)
# Route deterministic requests away from the LLM.
if route == "metrics":
return {
"answer": "Retrieve metrics directly from the monitoring system.",
"sources": [],
}
# Retrieve context once.
docs = index.similarity_search(question, k=4)
context = build_context(docs)
source_ids = [
doc.metadata.get("source")
for doc in docs
if doc.metadata.get("source")
]
# Check exact-match cache.
key = cache_key(question, source_ids)
if key in exact_match_cache:
return exact_match_cache[key]
# Select the most appropriate model.
model = choose_model(question)
# Keep trusted instructions separate from untrusted user input.
prompt_template = ChatPromptTemplate.from_messages(
[
(
"system",
(
"Answer the user's question using ONLY the provided context. "
"If the answer cannot be determined from the context, say so."
"\n\nContext:\n{context}"
),
),
("user", "{question}"),
]
)
chain = prompt_template | model
result = chain.invoke(
{
"context": context,
"question": question,
}
)
payload = validate_response(
answer=result.content,
sources=source_ids,
)
# Cache validated response.
exact_match_cache[key] = payload
return payload
except Exception:
logger.exception("Token-efficient pipeline failed.")
# Gracefully degrade instead of crashing.
return {
"answer": (
"The AI pipeline encountered an error. "
"Please continue using the standard operational workflow."
),
"sources": [],
}
What actually reduced token usage?
When teams instrument architectures like this, the largest savings rarely come from editing prompts. They come from eliminating unnecessary work.
The biggest improvements typically stem from:
- Retrieving documents once instead of multiple times.
- Caching semantically identical requests.
- Routing simple requests away from the LLM.
- Limiting context with explicit token budgets.
- Selecting the smallest suitable model.
These architectural shifts reduce cost and latency while making system behavior significantly easier to reason about.
Lessons learned
Several core principles consistently emerge when optimizing AI systems for production:
- Treat tokens like infrastructure: Tokens are a finite resource, just like CPU cycles or memory. Monitor them, budget them, and optimize them.
- Retrieval is usually the largest source of waste: Repeated retrieval often contributes more unnecessary tokens than verbose prompts. Share context whenever possible.
- Bigger models are not always better: Smaller, faster models effectively handle many operational tasks. Reserve larger models for genuinely complex reasoning.
- Caching is an engineering feature: A semantic cache is more than a performance optimization—it is a core architectural component that reduces cost, latency, and provider dependence.
- Measure before you optimize: Instrumentation must accompany every production deployment.
As AI systems mature, success will increasingly depend on engineering efficiency rather than raw model size. The hidden tax of AI agents is rarely a single expensive prompt; it is the accumulation of redundant retrievals, oversized contexts, unnecessary model calls, and repeated reasoning across distributed workflows.
“The most effective production AI systems are not the ones that generate the most tokens. They are the ones that generate only the tokens they truly need.”
By treating token consumption as a systems engineering problem, organizations can build AI platforms that are faster, less expensive, and highly scalable. Routing requests intelligently, budgeting context, sharing retrieval results, validating structured outputs, and introducing semantic caching are practical techniques that guarantee efficiency without compromising quality.
The most effective production AI systems are not the ones that generate the most tokens. They are the ones that generate only the tokens they truly need.
The post Stop the token bleed: building token-efficient multi-agent systems appeared first on The New Stack.
-
Robotics & Automation News
- How Data-Driven Quality Control Prevents Costly Defects in Modern Manufacturing