❌

Normal view

Received — 10 July 2026 ⏭ AI Infrastructure Archives - The New Stack

Meta’s Iris push signals the next phase of AI infrastructure

Abstract digital illustration of a circuit board pattern with interconnected nodes and pathways in cyan and black, representing technology infrastructure and connectivity.

Meta is preparing to manufacture its own AI chip for the first time. According to an internal memo, the company expects production of its proprietary processor, Iris, to begin in September.

After clearing bug testing in about six weeks, the chip — reported on by Reuters — is expected to take on some of the inference work currently running on third-party GPUs, giving Meta more control over how it builds and scales its AI infrastructure.

It’s unmistakable that this could be the company’s most important move yet toward in-house silicon for AI workloads.

But anyone can see this isn’t really about the hardware. It’s unmistakable that this could be the company’s most important move yet toward in-house silicon for AI workloads. The timing, as Meta is locked in an aggressive multi-billion-dollar infrastructure race, is critical. It’s clear that the company’s CEO, Mark Zuckerberg, wants to grow into the AI titan he believes the company can be, but it’s nearly impossible when the competition controls the core infrastructure.

Custom silicon for inference

Iris is designed for a specific job inside Meta’s AI infrastructure as custom silicon optimized for Meta’s heavy workloads.  Iris expands Meta’s Meta Training and Inference Accelerators (MTIA) program, which is intended to move targeted AI inference workloads onto custom silicon.

The processor would handle workloads that drive content ranking, recommendations, and generative AI services across Meta’s family of applications, including Facebook, Instagram, and WhatsApp.

  • The MTIA 300 is already deployed in production to run ranking and recommendation inference across Meta’s platforms.
  • The 450 and 500 variants target generative image and video inference through 2027.

By shifting these high-volume inference tasks to custom silicon, Meta can lower data center costs while bypassing the traditional hardware supply bottleneck for its day-to-day operations.

Securing the AI supply chain

Meta’s modular, rapid-fire approach to custom silicon is aggressive versus traditional industry timelines. The company plans to drop a new iteration roughly every six months through 2027.

Meta is working with Broadcom to design Iris, while TSMC will manufacture the chip. But custom silicon is only one piece of the equation. Scaling AI infrastructure also requires a steady supply of memory, storage, and networking components at a time when demand for AI hardware continues to strain global supply chains.

To support that expansion, Meta has also been securing key components across its supply chain. The company has signed long-term agreements for high-bandwidth memory from Samsung Electronics, flash storage from SanDisk, and fiber-optic networking equipment from Sumitomo Electric.

The strategy mirrors similar investments by other hyperscalers. Google continues to expand its TPU program, while Amazon has developed its Trainium and Inferentia processors.

Scaling to 14 gigawatts

The Iris rollout is one component of Meta’s broader AI infrastructure expansion. The company plans to bring roughly 7 gigawatts of computing capacity online this year, then double that to 14 gigawatts in 2027. At that scale, Meta’s AI infrastructure would consume more electricity than many small countries.

At that scale, Meta’s AI infrastructure would consume more electricity than many small countries.

And, scaling AI infrastructure at this level comes with an enormous price tag. Meta has projected 2026 capital expenditures of between $125 billion and $145 billion, making it one of the largest single-year infrastructure investors in corporate history

Meta just pulled off something rare, which is essentially convincing Wall Street that spending more money is actually a good thing.

Wall Street rewards AI spending

Yet Meta just pulled off something rare: essentially convincing Wall Street that spending more money is actually a good thing. Following a trillion-dollar wipeout in tech market cap amid investor nervousness about the sheer scale of AI spending, Meta’s shares climbed roughly 8%.

With new MTIA chips planned roughly every six months through 2027, Meta is betting that vertically integrated AI hardware can deliver lower inference costs and better performance than relying exclusively on merchant silicon. By bringing chip design in-house and securing critical components across its supply chain, the company is slated to scale AI infrastructure with greater control over cost, deployment, and optimization.

The post Meta’s Iris push signals the next phase of AI infrastructure appeared first on The New Stack.

Why retrieval quality is becoming the defining challenge in AI agent architecture

Neon digital waves and scattered data particles on a dark background, representing hybrid search, data pipelines, and AI engineering infrastructure.

Agentic systems usually have two jobs: Build context, then use that context to produce an answer or action.

Many failures that look like LLM problems start in the context-building step. The answer the LLM gives is limited by the context it was given, or it finds through tool calls. If the agent model cannot find the right sources, then improving the generation model will not improve the overall system.

“Many failures that look like LLM problems start in the context-building step.”

A client, Specstory, wanted to give users the ability to ask questions from the agent’s history. For example, why a team chose Authlib for authentication and what alternatives they considered. The chatbot needs the right prior conversations, decisions, and tradeoffs from a large corpus of coding sessions. The model and system prompt help only after those chat turns have been retrieved and are in context.

If retrieval ranks implementation snippets above the discussion where the team weighed alternatives, the agent can still produce a confident answer. It may find code that imports Authlib and a few inline comments, then describe the decision based on implementation evidence rather than the actual trade-off discussion.

The same pattern showed up in an AnkiHub operator review in our private community. A request for help studying based on lecture slides only works if the agent’s tool calls retrieve the right flashcards. The hard part is not finding any related cards. A lecture on the function of the heart may match hundreds of cards. Ranking decides whether the core cards make it into context or whether the system has to raise top_k and flood the prompt.

The exact setup changes by product. The context-building step might use local search, semantic search, web or API calls, or database queries. It might be handled by an agent, a fixed workflow, or application code. The process stays the same: gather the right context, then generate from it.

For example, a coding agent runs rg, opens files, reads logs, and inspects tests before writing a patch. A research agent searches the web and internal notes before writing an answer. A study assistant searches deck facts and user context before suggesting what to learn next.

When context building fails, the symptoms look like generation failures.

Retrieval failures mimic generation bugs

SymptomRetrieval cause
HallucinationThe answer source never made it into context.
Context rotLow recall forces a high top_k, so noisy results fill the context window.
LatencyWeak retrieval leads to more tool calls, larger candidate sets, and larger context windows.

A better model helps with reasoning and writing, but it cannot give a better answer without the right context.

“A better model helps with reasoning and writing, but it cannot give a better answer without the right context.”

The Mixedbread OfficeQA-Pro Eval shows the same pattern at the benchmark scale. OfficeQA-Pro uses 89,000 pages of financial documents, dense tables, scanned PDFs, and questions that require reasoning across documents. Giving Codex better search tools reduced tool calls and improved answer quality.

A scatter plot mapping Accuracy (%) against Tool Calls for three AI configurations.

Plain-text tools like grep and rg work (ish) on flat code files. They do not work well when context lives in PDFs, tables, chat histories, multi-modal inputs, web results, and permissioned data. In those cases, the agent needs a retrieval that can combine exact terms, meaning, metadata, permissions, and ranking quality.

Retrieval needs traces and evals

Once retrieval enters the architecture, the next question is whether it finds the right information.

For that, you need traces and evals. For each retrieval step, the minimum trace is the input, the outputs, and a way to label whether each output was relevant.

A flowchart diagram illustrating a data workflow where a horizontal sequence connects four steps: Input, Tool call, Output, and Label.

For a coding agent using rg the input is the command, the output is the returned snippets, and the label says which snippets helped, which were noise, and which relevant files were missing.

For product retrieval, the step might be BM25, semantic search, hybrid search with reranking, a generated SQL query, or something else. Capture the query or arguments, the returned documents or chunks, and whether those results were helpful.

Trace each retrieval step by itself, then evaluate the full context-building pass. The local trace answers “Did this query return useful material?” The full trace answers “Did the system collect everything the model needed before generation?” If it did, failures are a generation problem. If not, it’s a retrieval problem.

You cannot know where the failure started or what to fix without traces.

Different failures need different fixes

“Improve retrieval” is too broad to be useful, as different problems require different solutions. If a relevant document is missing, the trace should show where it disappeared: query building, retrieval, filtering, ranking, or final context assembly.

A sequential flowchart which maps a five-stage pipeline—Query builder, Retriever, Filters, Ranking, and Context—with each stage pointing down to its respective failure mode.

The failed step, plus what the trace shows, tells you what change to make.

Failed stepWhat the trace showsChange to make
rg / grepA conceptual query returns literal matches while missing relevant files.Add semantic search over files or chunks, or generate better keyword queries before calling rg.
BM25The query uses the right concept but different words from the source material.Add semantic search, synonyms, or query expansion.
Semantic searchExact names, error strings, document IDs, or domain terms are missing from the results.Add a keyword or BM25 path, or boost exact term matches.
Hybrid retrievalThe relevant passage is ranked 7th, but the context only takes the top 5.Add or tune a reranker, or raise candidate top_k before reranking.

The right fix depends on what the system was trying to retrieve. A decision-history question requires the decision, the alternatives, and the chats in which the team worked through them. A study question depends on the lecture material, deck metadata, semantic matches, and the user’s study context.

The architecture

Once you trace individual retrieval calls, the full architecture has a simple shape: fan out to context-building tools, then fan in to generate the final output.

A system architecture diagram showing a RAG pipeline.

The retrieval layer might be a search engine, a vector database, an SQL query, a local file tool, a web search API, or a custom service. The pattern stays the same: build candidate context, narrow it, rank it, assemble it, then generate from it.

Give agents human search controls

Semantic search compares embeddings (numerical representations of meaning). It helps when wording differs, but most retrieval intents also depend on structured constraints. A meeting search box can use semantic search over transcripts and notes, but a useful interface also lets someone filter by person, date, project, and source. 

A finance search may need the latest filing, a specific quarter, or an official source in addition to the closest semantic match. In e-commerce, the best semantic match for “32×30 cargo pants” may be an out-of-stock product. The system still has to decide whether to hide it, return it with a backorder note, or show it so the user can check later. That product decision is a retrieval decision because it changes which candidates reach the agent.

In a chat interface, those controls are in the tool schema, query planner, or app logic. If an agent runs the search, it needs arguments for the same constraints a human would set with filters, sliders, tabs, and sort menus.

A retrieval system usually needs several controls working together:

ControlWhat it doesExample
Exact matchMatches names, IDs, error strings, quoted phrases, tickers, or product codes.Find EADDRINUSE, Authlib, or a specific SEC accession number.
Semantic matchFinds related content when the wording differs.Find the meeting where the team discussed authentication tradeoffs.
Hard filtersRemoves invalid results before ranking.Limit by tenant, permissions, person, date range, size, or stock status.
SortsOrders candidates by a structured field.Prefer the newest, latest filing, lowest price, highest rating, or recency.
RankingScores candidates based on their likely usefulness for this request.Combine semantic match, exact match, freshness, source quality, and use.
RerankingUses a slower model or scorer on a smaller candidate set.Compare the query against the top 100 candidates before returning 10.

Here, a chunk means a small piece of source content, and a candidate is a chunk returned by the first search step. Ranking is the scoring step that orders those candidates. Context assembly then selects which chunks and structured fields to include in the model prompt.

Better ranking improves precision, which means a larger share of the returned chunks is useful. If the relevant chunks are near the top, the system can pass fewer chunks to the model, use fewer tokens, reduce latency, and expose the model to less noise. If the right chunk is ranked 40th and the context only includes the top 10, the system behaves as if the retrieval missed it.

People and agents use the same basic search path: ask for results, inspect what comes back, and decide what to use. A person can skim ten search results, compare titles, snippets, dates, domains, and URLs, and decide whether the result set looks right. They can open the third result, ignore the rest, and search again with a better query. An agent usually receives a bounded set of returned documents and reasons from the context. If the right source falls below the cutoff, the agent may answer from partial context. To avoid this, the system has to retrieve more candidates, run more searches, or pass more evidence into the model.

A missed document can change what the agent searches for next. Suppose someone asks why the team chose Authlib. If the first search misses the transcript where the team compared Authlib with alternatives, the agent may search the codebase instead. It finds imports, callback handlers, tests, and maybe a comment. Then it asks follow-up questions about OAuth configuration. The context starts to look complete, but it supports the wrong answer. It explains how Authlib was used and why it makes sense in the codebase, not why the team chose it.

“The context starts to look complete, but it supports the wrong answer.”

But ranking cannot repair every search problem. If the agent failed to request the latest filing, a reranker may faithfully select an older document with a closer wording match. If the tool has no date_range, person, source_type, size, or in_stock argument, the model has to impose hard constraints in the search text and hope that retrieval infers them. A hard filter gives the model less to infer, making semantic search more reliable.

Scale changes the retrieval problem

Search systems already have tools for this: indexes, filters, facets, sorts, caching, bounded reranking, and freshness jobs. Agent systems need the same discipline.

A human might search, adjust a date filter, scan the first page, then search again. One agent request can do that many times in seconds: rewrite the query, run keyword and semantic search, inspect thin results, issue follow-up searches, fetch sources for citations, and ask for more context before answering. With many concurrent users or agents, the retrieval layer can become a bottleneck.

Humans often wait through a slow search if the result is good. Agent systems often turn slow and uncertain search into more work. When ranking is weak, teams compensate by raising top_k, running keyword and semantic searches in parallel, adding reranking, fetching more source documents, and passing larger evidence bundles to the model. That can improve answers, but it moves the cost into tokens, latency, and retrieval load. A better ranking lets the system return fewer, better candidates, rather than making every request carry a larger pile of possible evidence.

With a small corpus, you can still search comprehensively quickly and cheaply, even with fully agentic approaches. That’s what I recommend when you’re starting and don’t have much data. Don’t add complexity until you need it. But with millions or billions of chunks, every extra retrieval call, candidate, ranking pass, and returned token adds up quickly.

Multi-stage retrieval is the production shape

Most production systems should split retrieval into stages, even when the UI is a chat box.

StageWhat happensTrace question
Search argument constructionThe app or agent turns the request and state into a query, filters, and sort.Did it ask for the right content with the right constraints?
Candidate generationThe system finds plausible chunks from text, vectors, or structured data.Did the right source enter the candidate set?
FilteringPermissions and product constraints narrow what can be returned.Was the source correctly excluded or wrongly lost?
SortingStructured fields order results when order matters.Was the latest, cheapest, highest-rated, or current item surfaced?
RankingThe system scores the candidates based on their usefulness for this request.Was the source present but ranked too low?
Summary returnThe system returns only the fields the agent needs.Did the app receive usable evidence and provenance?
Context assemblyThe app selects, formats, and budgets evidence for the model.Did useful evidence get dropped before generation?
EvaluationHumans or automated checks label whether the retrieval path worked.Can the team turn the failure into a specific fix?

Each stage leaves a different repair path. If the agent chose the wrong filters, changing the embedding model will not help. If the latest document was available but the tool never sorted by date, the fix belongs in the search arguments or retrieval API. If the right source was present but below the cutoff, the fix belongs in the ranking. If the right source came back but was dropped before generation, the bug is in context assembly.

As retrieval becomes a core part of agent architecture, teams increasingly need infrastructure that can combine semantic search, exact matching, filtering, ranking, and large-scale retrieval in a single system. Depending on requirements, this may involve search and retrieval platforms such as Vespa, Elastic, or Coveo, each of which supports different approaches to ranking, retrieval, and operational scale. 

The important point is not the specific technology choice, but recognizing that retrieval quality has become a first-class engineering concern. As agent workloads grow, retrieval systems are increasingly determining the accuracy, cost, latency, and reliability of the overall application.

The post Why retrieval quality is becoming the defining challenge in AI agent architecture appeared first on The New Stack.

OpenAI, Microsoft & Anthropic agree on who runs the agent. They disagree on what you can take back.

Colorful digital static resembling TV signal noise, evoking uncertainty over how AI agents like ChatGPT Work and Claude Cowork manage control, state, and data.

OpenAI announced ChatGPT Work on July 9 and began rolling it out to Pro, Enterprise, and Edu users. It runs on the new GPT-5.6, opens a user’s local files, edits Google Workspace and Microsoft 365 documents, and carries a multi-step task through to a finished deliverable.

Reuters placed it directly against Anthropic’s Claude Cowork, and both target the same person — a non-coder who wants the power of a coding agent without the terminal. Counting Anthropic, Microsoft, Perplexity, and Amazon, five leading labs have now released an agent of this kind. The batch shows that the newest agents are organized more by their intended users than by their functions.

Based on the target user persona, four archetypes appear: the knowledge worker, the power user who self-hosts, the developer, and the enterprise. The personas often overlap since one individual can embody all three roles. Additionally, a product like Claude Code caters to both solo developers and platform teams.

Therefore, consider these deployment archetypes categorized by the main buyer, rather than strict separations. The archetype only represents what marketing promotes. Behind the scenes, each lab has almost consistently decided who owns the runtime, persists memory, manages credentials, and enforces policy.

Four archetypes, based on the user persona

Let’s analyze each of the four archetypes individually, as each differs in the level of control available to users.

The first archetype serves the knowledge worker. A vendor operates the runtime and sells the agent as a delegation to someone who lives in documents rather than code.

ChatGPT Work is the newest, alongside Claude Cowork, which now runs cloud sessions on web and mobile while keeping local-file access on the desktop; Microsoft’s Copilot Cowork, a cloud-hosted agent that executes long-running tasks inside the Microsoft 365 trust boundary; Perplexity Computer, which works across local files and Microsoft apps; and Amazon Quick, the successor to Q Business as that product closes to new customers at the end of July. The user grants access and supervises the result. In most cases, the vendor manages the runtime and persisted state, except for Perplexity’s local option.

A second archetype belongs to the power user who self-hosts. The provider controls the persistent agent process and chooses where to store the state and credentials, often on a Mac mini that has become a piece of personal infrastructure in its own right.

OpenClaw and Hermes are the reference examples, open source, and run on the operator’s own machine. Self-hosting involves managing the control plane rather than full local custody, since both options still allow access to a hosted model and the storage of credentials for external services.

Related reads:

“Microsoft has proved it can survive major changes in the tides of technology… Today, it faces another evolution in one of its core cash cows, as late-stage unicorns and AI labs alike push deeper into Office territory.”

→  Read more in Cautious Optimism

The developer gets the third archetype, whose runtime spans the IDE, the terminal, the repository, and a cloud sandbox. Claude Code, OpenAI Codex, GitHub Copilot in agent mode, and the open-source OpenCode all live here, and Amazon’s developer agent is folding into its Kiro tool. Coding-agent execution is extending from the local IDE to vendor-managed sandboxes and asynchronous cloud workers, making this the most challenging archetype to categorize clearly.

The fourth archetype is built for enterprise workflows and integration with business processes. They run an open agent framework, such as LangGraph or CrewAI, on a managed, governed runtime. ADK on the Gemini Enterprise Agent Platform, Strands on the Bedrock AgentCore, Microsoft Agent Framework on Foundry Agent Service, and Claude Managed Agents belong to this category. OpenAI’s Agents SDK can be hosted on some of these runtimes, including AgentCore, which AWS lists as one of its supported frameworks. The vendor operates the infrastructure, and the customer configures identity, policy, and retention on top of it.

PersonaRepresentative productsRuntime ownershipState and credentialsPlatform type
Knowledge workerChatGPT Work, Claude Cowork, Copilot Cowork, Perplexity Computer, Amazon QuickVendor cloud for most, with per-folder local access on someVendor persists session state, user grants scoped credentialsPackaged experience
Power user, self-hostOpenClaw, HermesThe operator’s own machineOperator chooses where state and tokens live, though inference is often externalPackaged experience, self-operated
DeveloperClaude Code, OpenAI Codex, GitHub Copilot, OpenCodeSplit across the IDE, the laptop, and a cloud sandboxRepo and local for now, drifting into hosted sandboxesSwing, moving toward platform
Enterprise, workflow-drivenADK on Agent Platform, Strands on AgentCore, MAF on Foundry, Claude Managed AgentsManaged vendor runtime, customer-configurableCustomer defines identity, policy, and retention; platform brokersProgrammable platform

The line below the personas

The persona-based approach abstracts four things: where execution runs, where state is persisted, how authority is delegated, and where policy is enforced.

Anthropic describes its design as decoupling the brain from the hands. The harness that calls Claude runs separately from the sandbox where code executes, and a session, an append-only log of every model call, tool call, and result, connects the two. Because the sandbox is kept separate from the brain, the agent can start reasoning before any container exists, and the code it runs remains far from the developer’s credentials.

The same four planes show up at the other vendors. AgentCore Runtime gives each session a dedicated microVM with an isolated CPU, memory, and filesystem, and meters compute usage. Google can route governed traffic through its Agent Gateway, where Model Armor policies inspect configured ingress and egress flows, while Agent Identity and an Agent Registry track the fleet. Microsoft assigns each hosted agent a dedicated Entra Agent ID and runs it in a per-session sandbox whose filesystem survives idle periods.

Deploying agents on a managed runtime is more like leasing a workshop than purchasing a tool… a long-term tenant installs their own locks, maintains their records, and takes their tools when the lease ends.

Deploying agents on a managed runtime is more like leasing a workshop than purchasing a tool. The landlord manages the building and supplies the power, but a long-term tenant installs their own locks, maintains their records, and takes their tools when the lease ends. The personas often conceal this difference. Currently, the vendor typically operates the runtime, so the key question is how much control the customer can still exert and what they can take away across the four planes.

Where the line falls

What differentiates each offering is not the compute operator, since vendors handle nearly all of it. It is how much of those four planes a product leaves the customer to configure and export. A packaged experience hands nearly all four to the vendor and returns supervision and a finished outcome. A programmable platform operates the infrastructure but lets the customer define identity, policy, and retention and move the code elsewhere.

Copilot Cowork shows that the two axes are separate. It is a packaged knowledge-worker experience, yet it runs on a governed enterprise platform and inherits Microsoft’s identity, compliance, and audit controls.

The persona sells the product, but the four planes decide the lock-in.

A product can be packaged on the surface and programmable underneath, which is why the personas and the control planes have to be read as different questions. ChatGPT Work makes the same point from the developer side, since OpenAI’s new desktop app folds Chat, Work, and Codex into a single surface, though OpenAI has not detailed how far the runtime or credential store are shared beneath it. The persona sells the product, but the four planes decide the lock-in.

UsecaseAgent TypeTradeoff
Delegate a knowledge task to an agent you superviseA knowledge-worker agent such as ChatGPT Work, Copilot Cowork, or Amazon QuickThe vendor typically operates the runtime and persists state, and you configure little below the surface beyond access and approval
Keep state and credentials on hardware you controlA self-hosted agent such as OpenClaw or Hermes, in a local configurationYou control the persistent process, though model inference and some tools may still be remote
Ship code changes across the IDE, repo, and CIA developer coding agent such as Claude Code, Codex, or CopilotExecution spans your tools and a cloud sandbox, so ownership is split and worth mapping before you commit
Run many governed agents with audit and identityAn enterprise runtime platform such as AgentCore, Agent Platform, Foundry, or Managed AgentsThe vendor operates the infrastructure while you define identity, policy, and retention, in exchange for coupling workflow logic to one cloud

Real deployments combine the rows rather than picking one. Teams on Foundry Agent Service commonly run open-source orchestration, such as LangGraph, for agent logic while leaning on the platform for governed execution, and Microsoft’s own hosted runtime now supports long-running personal agents like OpenClaw and Hermes with durable state. The boundary between experience and platform is not a thick, well-defined boundary, but a thin line a single system can cross, bridging rival camps.

The agent market is not being decided by open against closed. Open frameworks like Strands, ADK, and Microsoft Agent Framework are precisely what the governed runtimes are built to host.

The agent market is not being decided by open against closed. Open frameworks like Strands, ADK, and Microsoft Agent Framework are precisely what the governed runtimes are built to host. Vendors now manage the runtime across nearly all archetypes, shifting the competition to who can most effectively configure and export state, identity, and underlying policy.

If coding agents enter managed sandboxes alongside knowledge-worker agents, the developer archetype will be established on the platform side. The map will then transform into what the vendors are already outlining. When enterprise teams assess agents, their most important question should be how much of execution, state, identity, and policy they can configure and extract from the product. They should not focus on which persona it presents or whether its framework is open, as the framework no longer determines the product’s value or locking-in capabilities.

The post OpenAI, Microsoft & Anthropic agree on who runs the agent. They disagree on what you can take back. appeared first on The New Stack.

❌