❌

Normal view

Perplexity’s AI agents helped build a database. They weren’t allowed to run it.

Abstract glitch wave

Perplexity decided it was paying too much for DynamoDB and wasn’t getting the control it wanted over read performance. So it built its own database: CobbleDB.

Built by two engineers in two months with help from hundreds of persistent coding agents throughout development, CobbleDB is a roughly 40,000-line Rust key-value store that now handles part of Perplexity’s production search traffic. The company measured median batch-read latency at 5.6 milliseconds after the move, compared with 31.4 ms on DynamoDB before the cutover, while p99 went from 123 ms to 24.2 ms.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

But the database itself is only part of the story. CMU professor Andy Pavlo argued at Percona Live earlier this year that databases are the hardest and most important challenge for AI agents, in part because mistakes involving production data can be difficult or impossible to reverse.

Perplexity went ahead and used hundreds of agents to help build one anyway, but they weren’t given the keys to production.

It’s expected to cost at least 20% less than DynamoDB and plans to open-source the database at some point.

Why DynamoDB couldn’t keep up

Each search requires the serving layer to retrieve pre-chunked passages and vector embeddings, with a single Search API call fetching 100 to 120 page keys in batches of 10 to 20. Each item averages about 50 KB.

DynamoDB gave Perplexity little control over how it handled reads, which meant a slow replica could hold up the entire things. It also charged for the steady flow of large reads and writes generated by search, crawling, and reprocessing, which made cloud costs difficult to justify as traffic and the corpus grew.

That led Perplexity to separate long-term document storage from the database serving live searches.

Three tiers for search data

The storage stack is split into three pieces. Pillar keeps durable document state in YTsaurus on HDDs, including versioned metadata, chunks and embeddings, while Lorry packages updates into partition-specific batches and moves them through S3 to CobbleDB.

Processed page data is spread across three replicas per partition, with hashed URLs as keys and RocksDB keeping often accessed data in memory while the rest stays on local NVMe. Reads stay within the same availability zone when possible, and the router can try another replica if one is slow rather than hold up the batch.

Updates come through S3 and are applied independently, allowing a replica to fall behind and catch up without blocking the others.

Roughly 5X Lower Batch-Read Latency

Perplexity was handling approximately 200,000 requests per second when it measured CobbleDB at 5.6 ms for a median batch read, down from the 31.4 ms it had recorded on DynamoDB. At p99, latency went from 123 ms to 24.2 ms.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

The comparison comes with an important caveat; DynamoDB and CobbleDB weren’t tested side by side against identical traffic: the DynamoDB figures were recorded before the cutover, and CobbleDB’s afterward. Perplexity separately ran synthetic benchmarks using batches of 10 to 15 keys with values ranging from 100 bytes to 100 KiB.

Its cost model puts CobbleDB at least 20% below DynamoDB across the commitment options evaluated, though that estimate doesn’t include the engineering cost of supporting the database.

In later load testing, CobbleDB reached 500,000 requests per second before performance started to decline.

Agents built it, engineers controlled it

The agents carried context across sessions, catching problems with restore assumptions and runtime configuration while working on fixes and tests. But they weren’t running the database.

The two engineers kept control of the architecture and production system, particularly important given Pavlo’s warning about putting agents near critical production data.

Ownership has long-term costs

Shipping CobbleDB in eight weeks solved Perplexity’s immediate engineering bottleneck, but maintaining a custom datastore could prove considerably harder. The latency results aren’t from a controlled side-by-side benchmark, and the projected savings don’t include the engineers needed to maintain CobbleDB and respond when something breaks.

Like Shopify and Ramp, which built custom coding agents around third-party models, Perplexity kept the cloud infrastructure but replaced a managed service with something built for its own needs. CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

CobbleDB shows how AI-assisted development is changing that calculation, making custom infrastructure more practical for smaller engineering teams.

The post Perplexity’s AI agents helped build a database. They weren’t allowed to run it. appeared first on The New Stack.

Anthropic bet users were choosing wrong. So it removed the choice.

Single lane

Using Claude for anything beyond a quick question has always started with a routing decision to use Chat or Cowork? Anthropic has decided to eliminate that fork.

Starting Wednesday, Claude Chat and Cowork merge into a single interface where one conversation can handle everything from a simple answer to a multi-step project with connected tools and background execution. The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design — previously a standalone workspace — into conversations.

The combined effect promises a streamlined experience with Claude picking up context, skills, and connectors as the work requires, and can keep running after you close your laptop.

The company is also launching Claude Docs and Claude Slides in beta on paid plans, and moving Claude Design, previously a standalone workspace, into conversations.

Two modes, one problem

Anthropic built Cowork as a desktop-first agent for bigger work and Design as a separate workspace for visual output. Both shipped earlier this year and gained traction.

“We built Cowork as a separate place for bigger work, and Design for visual work,” Anthropic said in its announcement. “People used both, and told us the frustrating part was deciding where a task belonged.”

Anthropic has run into this problem before. When the company promised 20x more usage on its Max plan, developers complained that it wasn’t always clear where one limit ended and another began. Cowork and Design created a similar headache by making people decide where to start the work before they could actually start it.

Context didn’t always follow the work either, so moving from chat to Cowork or Design could mean bringing the same background along all over again. Anthropic addressed part of this in August when it unified Claude’s memory across chat and Cowork. Wednesday’s change goes further by merging the products themselves.

“We built Cowork as a separate place for bigger work, and Design for visual work,”

Context that finally travels

Cowork’s capabilities — local file access, multi-step execution, scheduled tasks and connected tools — now live inside the conversation. A workflow like that previously meant switching from chat to Cowork and carrying the context with it, and until Anthropic brought Cowork to web and mobile in July, it also required the desktop app.

Claude still asks before taking an action by default, but it can be set to keep working and check in only when something needs a closer look, while recurring tasks such as a weekly report can be scheduled to run every Monday without being started manually.

Output stays in-conversation

Claude Docs and Claude Slides launch in beta on paid plans, bringing document editing and presentation building directly into the app. Claude can turn work from an existing conversation into slides, which can then be edited, presented from Claude or downloaded as PowerPoint or PDF files.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice. Everything stays attached to the conversation that produced it.

Claude Design also now works inside conversations, in addition to remaining available on its own. For organizations that rely on MCP connectors to wire Claude into external tools and data, the merge means those connections are available wherever a conversation goes — without requiring users to start in a specific mode. Skills, connectors, and artifacts carry across what used to be product boundaries.

The practical benefit is that a report and a slide deck based on it don’t have to begin as separate jobs with the same background supplied twice.

What Anthropic hasn’t said

The announcement leaves some gaps. It doesn’t say whether users can force a request to stay in simple chat mode rather than letting Claude decide how to handle it, or how that decision affects context windows and token consumption. There’s no mention of API changes, which makes this a consumer and team product shift, not a platform one, at least for now.

It also doesn’t address what happens to workflows built around the old separation. Shopify rebuilt its mobile development stack in 12 weeks when it consolidated tools that had grown apart — the question for Claude power users is whether their existing Cowork setups, skills, and scheduled tasks survive the merge cleanly. Anthropic says existing Cowork chats, projects, artifacts, connectors, and skills will remain available.

Rollout starts with Pro

The unified interface rolls out to Pro and Max users across web, desktop, and mobile over the next few weeks. Anthropic says there’s nothing to enable. Team and Free plans follow. Enterprise customers are on a separate timeline; Anthropic is giving administrators at least 30 days’ notice before the change reaches their organizations.

The post Anthropic bet users were choosing wrong. So it removed the choice. appeared first on The New Stack.

AI evaluator: The most important AI job in history? How developers might fill the proposed new job

Lots of pink escape keys

The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for changes in how the industry is regulated and develops. In a three-part plan that includes both democratic and global coordination, Amodei writes that the first step was something Anthropic is committing to unilaterally.

“Each frontier AI company [should] commit to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes,” writes Amodei.

Amodei’s essay followed dire warnings from former Anthropic and OpenAI pretraining research specialist Jacob Coxon, who posted a thread on X saying the people building AI earnestly “believe that it could kill us all” by the end of the decade.

Shortly after Amodei published his essay, OpenAI CEO Sam Altman and SpaceXAI founder Elon Musk chimed in: “I agree with Dario,” posted Altman; “Dario is right,” posted Musk. Later that day, Demis Hassabis, founder of Google DeepMind, posted, “Dario’s essay points towards the right path forward.” In a post on X, Meta CEO Mark Zuckerberg writes that Meta Superintelligence Labs already uses independent evaluators, and that, “In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.”

This week, theories began to surface about why the world’s biggest frontier AI labs would want to intentionally slow their pace when competition is so fierce. “The desire to slow down is puzzling, but perhaps if the whole system slows down, the rules of winning can be the same for all,” posted Nikesh Arora, chairman and CEO of Palo Alto Networks.

In his essay, Amodei likens the proposed job of independent AI evaluator to embedded regulatory supervisors used in the banking industry, i.e., third-party professionals. Altman describes the job as having “employee-like access” in his post on X.

So, who could fill these roles that AI leaders agree are desperately needed?

Salaries top out at $687K; are you interested?

METR’s current job openings are well paid (salaries top out at around $687,000), and the job specs are daunting. 

“You’re scrappy, creative, independent, and self-directed (because during the exercises you’ll only have a few other METR employees you can talk to). The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases,” reads the spec.

“You’re scrappy, creative, independent, and self-directed. The work is novel, and you’ll need to figure a lot of stuff out on the fly largely by yourself. You are excellent at loss-of-control threat modeling and breaking down safety cases.”

Similar but less colorfully illustrated roles (paid between $180K–$300K) are also available at AI model training company Mercor, where candidates will need a Ph.D. or M.S. and more than two years of work experience in a computer science, electrical engineering, econometrics, or another STEM field that provides a solid understanding of machine learning and model evaluation.

“Employee-like access fluctuates wildly”

AI security consultant and CTO at Komodo, Kadan Stadelmann, tells The New Stack that his typical week sees him work differently with each client. This is because “employee-like access fluctuates wildly”, from rigid focus areas to broad access, and much of that aspect is determined by contracts signed before work begins.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions,” Stadelmann says. “Among my duties are tasks that include monitoring chains-of-thought and prompts. The goal is to establish an objective and look at a specific AI system to question how autonomous the system is, and how long it takes to complete specific tasks. Most importantly, evaluators at my level monitor for how well a team adheres to its claimed safety practices.” 

Software engineering skills beat doctorates

Although METR wants evaluators to have a Ph.D. up their sleeve, Stadelmann says that as the prevalence of this role expands, he feels the technology industry has been, and continues to be, driven by people who can demonstrate strong engineering skills, not doctorates.

“I am invited into labs to probe numerous risk vectors, including agent behaviors under realistic conditions. Among my duties are tasks that include monitoring chains-of-thought and prompts.”

Questioned on whether costs create a barrier for smaller labs, Stadelmann notes that some AI evaluation work is funded by third-party non-profits, which protects independence. 

“But overall, evaluators will not be able to keep up with big frontier model firms. They will be out of control, and we will be dependent upon their own internal ethics. Plus, anyway, many of the smaller labs of any worth may inevitably be acquired by the big players in this space,” he adds.

What happens when an evaluator finds something wrong?

Founder and CEO of facial image AI identity governance company Indie Me, Dion Johnson, tells The New Stack that what interests him most about embedded AI evaluators isn’t the job title; it’s what happens when their judgment uncovers that the model behaved in a way nobody expected. 

“If the evaluator can only raise concerns when those concerns are convenient, then we have not created independent oversight — we have created another layer of process,” Johnson says. “The evaluator needs enough access to see the uncomfortable things, not just the polished demonstrations. They need to understand what happened during training, what failed during testing, what behaviors appeared unexpectedly, and where the team itself still has uncertainty.”

On the required skills AI evaluators need, Johnson agrees that technical depth, machine learning environment security experience, and software engineering as a whole matter.

“To choose a competent AI evaluator, I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not. This is someone who can sit in a room full of brilliant people at an AI model company on launch day and say ‘I’m not convinced’… and that takes judgment and courage,” he adds.

“I would look for someone who is deeply comfortable with uncertainty and deeply uncomfortable pretending they know something they do not.”

Fear and loathing in the AI space

In his September 8 post, which has been viewed 172 million times and seemingly spurred AI leaders to change course, Coxon, the former AI researcher at OpenAI and later Anthropic, writes: “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible but I hear the same people express fear privately. No other human activity poses this level of danger.”

As for where Coxon looks for work next, perhaps it might be a role in AI evaluation execution engineering.

The post AI evaluator: The most important AI job in history? How developers might fill the proposed new job appeared first on The New Stack.

❌