❌

Normal view

Received — 24 June 2026 ⏭ AI Infrastructure Archives - The New Stack

OpenAI wants to claim more of the AI stack with Jalapeño, its first custom chip

OpenAI on Wednesday announced Jalapeño, its first custom inference accelerator, co-developed with Broadcom and supported by Canadian electronics manufacturer Celestica, and the first step in its multi-generation compute platform. 

The AI company says Jalapeño was designed to work with all large language models (LLMs) and will help make AI faster, better, and cheaper. Behind that rosy mission, OpenAI isn’t shy about its desire to own the full AI stack, something more AI giants are already leaning into.  

“Those serious about platforms should be serious about silicon.”  

As Ben Bajarin, CEO and principal analyst at consumer technology research firm Creative Strategies, posted on X: “Those serious about platforms should be serious about silicon.”  

Remember my mantra, yes this is an intentional play on Alan Kay's quote.

Those serious about platforms should be serious about silicon. https://t.co/ghWmCqkCKe

— Ben Bajarin (@BenBajarin) June 24, 2026

But with few technical details released, developers are left wondering if OpenAI’s widening footprint will be empowering or restrictive. 

Get in, Big Tech. We’re all building in-house chips now. 

OpenAI isn’t the only Big Tech name to mint its own AI chips. 

Way back in 2016, Google designed and built its own custom hardware for TensorFlow, its machine learning software, the Tensor Processing Unit (TPU). A couple of years later, Amazon debuted AWS Inferentia, its first purpose-built chip for AI and ML. Trainium then hit the scene in 2022, shortly followed by Microsoft’s Azure Maia AI Accelerator in 2023. And nothing is certain yet, but in April, Reuters reported Anthropic is contemplating designing its own chips, though the AI company remains noncommittal for now, at least publicly. 

Why is everyone jumping on the custom-chip bandwagon? 

Blame the compute gold rush, as AI companies increasingly clamor for compute power — ”a compute-powered economy,” as Greg Brockman, president, chairman, and co-founder, OpenAI, puts it. And the numbers surely back it; Stanford’s 2025 AI Index Report says, “training compute doubles every five months.”

While building custom AI chips in-house doesn’t completely alleviate compute pressures, it is one way for OpenAI and its Big Tech brethren to expand compute capacity while potentially lowering costs and reducing reliance on third-party suppliers. 

Exciting claims, no proof

In its announcement blog post, OpenAI describes its new chip as “designed to be the best inference platform for LLMs.”

Specifically, Richard Ho, head of hardware at OpenAI, states: 

“We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits.”

But the AI company remains tight-lipped on any real technical details. 

While it claims current tests put Jalapeño’s performance “substantially better than current state-of-the-art,” it doesn’t provide benchmarks to back that up. Instead, it tells developers to expect a detailed technical report “in the coming months.”

What OpenAI does divulge is that engineering samples of the chip are currently running on ML workloads in its lab, including GPT-5.3-Codex-Spark.

Will Jalapeño serve developers, or is OpenAI’s desire to own the AI stack? 

OpenAI makes no qualms about its quest for full-stack control. In doing so, the AI company claims it will make its models “faster, more reliable, and more affordable for users.”

Its logic goes a little something like this: Better infrastructure means more efficient compute, which means better training, which means better models, which means better products, which means more revenue. Then, it explains, it can reinvest that revenue in its infrastructure to make intelligence better for everyone.

But given how little the AI company has revealed about the chip’s specs, it seems developers will have to sit back and watch where the chips fall. 

Jalapeño, then, is simply the next move in OpenAI’s quest to control the whole AI chessboard, moving beyond models and products to the underlying infrastructure itself. 

For developers, OpenAI seems adamant on insisting its full-stack strategy will lead to better performance and pricing for everyone and ultimately empower “anyone trying to learn, create, or solve hard problems.” Still, it’s worth considering: As OpenAI’s grip tightens, will developers become beholden to its ecosystem? 

Several times in its announcement, OpenAI reiterates that it designed Jalapeño for current and future LLMs — all of them. But given how little the AI company has revealed about the chip’s specs, it seems developers will have to sit back and watch where the chips fall. 

Built fast with a long roadmap ahead

The few behind-the-scenes details OpenAI does choose to share boast about its development speed, stating it brought Jalapeño from design to manufacturing tape-out in nine months — “what we believe to be the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors.”

The AI company chalks up that fast timeline, in part, to its own models accelerating parts of the design and optimization processes. 

Looking ahead, Jalapeño is slated for deployment at a gigawatt scale in Microsoft’s and other partners’ data centers by the end of the year. 

That’s just the beginning. OpenAI hints at an upcoming multi-generation roadmap, posing the question: What will it seek to control next? 

The post OpenAI wants to claim more of the AI stack with Jalapeño, its first custom chip appeared first on The New Stack.

Agentic infrastructure operations begin with accurate, reliable infrastructure data

An abstract, high-angle view of glowing orange and white light trails stretching across a dark, grid-patterned metallic surface, evoking futuristic infrastructure or high-speed data transmission.

Organizations are racing to apply AI across the enterprise, and infrastructure is one of the most compelling targets: automated provisioning, self-healing networks, and agents that deploy and manage servers without human intervention. The promise is real, but so is the risk. 

No matter the domain, AI agents are only as good as the data they’re given. Agents without a complete and accurate picture of the network and associated infrastructure will make confident mistakes. In infrastructure, those mistakes have brand and revenue-related consequences: exposed databases with PII, failed deployments, and outages that take the entire business offline. 

“Agents without a complete and accurate picture of the network and associated infrastructure will make confident mistakes.”

Most enterprise infrastructure is managed through a patchwork of siloed, fragmented tools: separate systems for IP address management, data center inventory, and device configuration. The list goes on.

Each system captures a slice of the picture, but none of them complete the full vision. HyperFRAME found that over 70% of industry leaders identified this as a core bottleneck. Agentic automation cannot solve this issue, but it will expose it through its failures.

Before you can trust an AI agent with your infrastructure, you need to give it something to trust: a single, unified model of what’s on your network, how it’s configured, and how it’s supposed to behave. According to NetBox Labs CEO and cofounder Kris Beevers, that’s an Infrastructure Intelligence platform.

What is infrastructure intelligence?

Whether run by AI or human agents, infrastructure is impossible to manage when critical systems contain unknowns. Infrastructure intelligence is the foundational blueprint of your infrastructure: a unified, continuously updated model that captures not just what exists, but what is intended, what has changed, and what needs attention. It is the prerequisite for automation at any scale.

“AI is raising the stakes for infrastructure management, and the challenge is no longer just documenting infrastructure; it’s also understanding it…”

“AI is raising the stakes for infrastructure management, and the challenge is no longer just documenting infrastructure; it’s also understanding it,” says Beevers. “A source of truth was enough for the last decade. But today, teams need context – a trusted, continuously updated understanding of infrastructure that helps them (and their AI agents) model, see, act, and govern with confidence. AI doesn’t eliminate the need for infrastructure data. It makes it more important than ever.” 

It starts with a system of record. More than just an inventory list: it is a living representation of the intended state (what everything is supposed to look like) and the operational state (what it actually looks like right now) of your network. The gap between these two states is drift, and that is where risk lives. Without a system that tracks both states simultaneously, your team is always reacting, chasing down misconfigurations, and manually reconciling tool outputs (hoping nothing critical slips through).

Full infrastructure context connects the intent and design to the operational state, providing drift detection, observability, and lifecycle management tools in a single continuous data thread. Instead of switching between five different tools to answer a single question about a specific device, your team and your agents have all the information they need in one place. What is the device supposed to be doing? What is it actually doing? When did it change, and who changed it? Full context means that these questions have immediate answers.

Guardrails close the loop. Both humans and AI agents can make well-intentioned errors, and in infrastructure, the blast radius of these errors can be severe. For this reason, your infrastructure must have well-defined audit trails, branching workflows, change management processes, and operational validation from the beginning, not as an afterthought after something goes wrong. 

Teams move from handholding every agent action to trusting the system to catch bad outcomes. Cautious early adoption quickly grows into confident, autonomous, scaled deployments.

The foundation for any automation journey

Agentic automation/Agentic NetOps is coming to infrastructure teams, whether they are ready or not.

No matter where a company is in its automation journey, Infrastructure Intelligence provides a strong foundation for everything else. Organizations that are early in their automation strategy have manual processes they want to automate — they need a clear picture of the environment to do this safely. Teams that are running agentic workflows across complex, multi-site networks share the same requirement: Infrastructure intelligence.

NetBox Labs, the commercial steward of the open-source NetBox, recently expanded its platform to ensure that every infrastructure management workflow can be addressed by agents. The announcements make infrastructure AI Agent-Native: Extending the NetBox MCP server across the entire NetBox Labs Platform and releasing an array of pre-built agent skills.

These agentic tools are designed to leverage the existing infrastructure intelligence from NetBox Labs’ systems, ensuring that all agentic network provisioning capabilities are combined with the required guardrails, validations, and protections to keep the network running smoothly. 

Adding agentic features across the entire NetBox Labs infrastructure intelligence platform gives agents unprecedented knowledge, skills, and power. Agents can access NetBox Data Exchange — the world’s largest database of infrastructure metadata. NetBox Assurance and Discovery helps teams and their agents identify and mitigate drift.

“Giving AI agents access to production infrastructure without guardrails is a recipe for outages.”

According to NetBox CEO and cofounder Kris Beevers, “The future isn’t just autonomous infrastructure. It’s a trustworthy infrastructure. We know that trust and governance are the foundation of AI-driven operations. Giving AI agents access to production infrastructure without guardrails is a recipe for outages. That’s why we’ve paired these AI agent native updates with new validation tools so teams can ask, “‘Is this change safe to deploy?’ and ‘What breaks if this fails?’”

The new validation tools help agents self-correct, validate changes, and meet compliance requirements, ensuring continuous compliance and pre-change safety within the System of Record.

AIOps teams that establish a foundation of infrastructure intelligence gain more than efficiency, visibility, and control. They gain the confidence to automate services in production without losing sleep over it. Agents stop guessing and operate from verified real-time data. Teams stop reacting and focus on building. And the organization does not see AI as a liability, but as a capability that can be expanded.

NetBox Lab’s new infrastructure intelligence platform is designed for both humans and agents, making it easier to manage infrastructure across every lifecycle stage — from design through end-of-life.

Whether you’re a NetBox open source user or NetBox Labs customer, you can celebrate NetBox turning 10 at its inaugural conference, NetBox Evolve, which will be in Florida at the Kennedy Space Center on October 13, 2026.

The post Agentic infrastructure operations begin with accurate, reliable infrastructure data appeared first on The New Stack.

Sakana Fugu is more than a router. But it’s not the blueprint for AI sovereignty, either.

This week, Sakana AI released Fugu, a multi-agent orchestration system designed to deliver frontier-model performance all while reducing the risks of relying on a single provider. 

The Japanese AI R&D company says Fugu performs as well as Anthropic’s Fable 5 and Mythos Preview on engineering, scientific, and reasoning benchmarks by breaking up tasks into subtasks and strategically routing them across a swappable pool of expert agents. But early reactions are mixed.

While Sakana positions Fugu’s “collective intelligence” as the blueprint for AI sovereignty, not all users report frontier-model-level performance. Others note fast burn rates and unnecessarily high prices. 

Many agree that, though interesting, Fugu likely won’t be the hero to AI sovereignty it hopes to be. 

Is this just another router? Not really. 

Sakana says Fugu’s internal routing logic is founded on its own research in learned model orchestration, specifically noting two papers, Trinity and the Conductor. 

Unlike multi-model routers, such as OpenRouter Fusion, that send a prompt to multiple models and then compare or combine the results, Fugu breaks down user prompts into subtasks and determines which subtask to send to which model. In this way, Fugu “dynamically orchestrates the world’s best models to tackle complex, multi-step tasks,” so Sakana says. 

From the outside, you just see what looks like one model, accessible via a single OpenAI-compatible API.

But what the AI company doesn’t tell you is how it decides which tasks get routed where; that information is proprietary. From the outside, you just see what looks like one model, accessible via a single OpenAI-compatible API.

“relying on a single company’s model for national infrastructure is a massive risk. As recent export controls have shown, access to top models can disappear overnight.”

Fugu doesn’t have to farm out every task, though. It’s a language model itself, specialized for model selection, delegation, verification, and synthesis internally, so it can also solve requests directly when its own response is sufficient.

A hero for AI sovereignty, it appears not

In an X post, Sakana CEO and co-founder David Ha writes, “relying on a single company’s model for national infrastructure is a massive risk. As recent export controls have shown, access to top models can disappear overnight.”

Human intelligence is fundamentally a collective intelligence. We solve complex problems by participating in a vast cultural network that builds upon ideas across generations.

I believe the strongest AI systems will become a collective intelligence, too.

Since we started Sakana… https://t.co/yulKqdei2c

— hardmaru (@hardmaru) June 22, 2026

That “massive risk” comment is likely a jab at what happened to Anthropic, when an export control directive forced the AI company to pull Fable 5 and Mythos 5 just three days after launch. 

See also: Fable 5 ban: 4 open models responded before Anthropic could restore access

Following this news, Sakana positions Fugu as the antidote to single-provider reliance. Because it relies on a pool of “entirely swappable agents,” the idea is that Fugu is less likely to leave users in a bind if one provider suddenly restricts access. It can simply route work to other models. 

The AI company considers this capability enough license to claim it’s “delivering the realistic, resilient blueprint required for AI sovereignty.” But some initial reactions call that hyperbolic: 

“This is just a highly advanced router/wrapper, not a fundamental leap like Mythos/Fable was,” argues one Redditor.

Though it’s likely not fair to call Fugu a simple multi-model router, its ultimate reliance on other models means it’s not the hero for AI sovereignty it aspires to be. After all, if more than one model provider restricts access at the same time, Fugu’s capabilities also take a hit. 

As another user writes on HackerNews: “As a developer outside the US I think it’s vital to have alternatives to OpenAI and Anthropic, but sadly this is not it,” calling out what they describe as the tool’s unfortunate price-to-burn-rate ratio, an “extremely slow” API, and poor quality in comparison to Fable:

“It’s nowhere remotely near usable as a day-to-day workhorse.”

Not all user reviews back up the benchmarks

Sakana points to coding, reasoning, science, and agent benchmarks to prove Fugu’s value, stating its tool consistently beats Gemini 3.1, Opus 4.8, and GPT 5.5.

Source: Sakana AI

It also highlights what it says is the success of its beta program, where almost 500 early users tested Fugu on lengthy, multi-step computational workflows.

In particular, it claims that one cybersecurity engineer confirmed Fugu successfully operated within parameters and avoided destructive actions, while other teams praised Fugu Ultra for besting GPT 5.5 in code review and maintaining an “unusually strong persona stability across long sessions.”

But moving from benchmarks and PR-ready examples to early community sentiment adds more color to the story. 

One user on HackerNews calls Fugu “quite strong” for a few agentic coding tasks, but notes they weren’t able to do many deep reviews before their quota ran out, adding: “For implementation I found it weaker, it made a few mistakes that I haven’t seen frontier models make in a long time.” 

A Redditor had a different experience. They, too, bemoan burn rate issues, but note: “It caught things Opus 4.8 ultra and codex 5.5xhigh clearly missed in a fairly large data ingestion / processing project.”

Some users question the price tag

Furu is generally available today in most regions (save the EU) in two tiers: a low-latency model that integrates with chatbots and tools like Codex for daily tasks and Fugu Ultra, the heavy-hitter that coordinates a deeper pool of experts for more complex, high-stakes tasks. (This is the one that’s supposed to rival Fable 5 and Mythos Preview.)

Subscription plans are available at $20, $100, and $200 monthly rates for both Fugu and Fugu Ultra. Pay-as-you-go pricing is also available, with Fugu billed at standard rates per underlying model, and Fugu Ultra running at $5 per million input tokens and $30 per million output tokens, with higher rates when context exceeds 272k.

Several early users on Reddit and HackerNews deem these price tags too high, especially when they’re experiencing what now feels like the soundtrack of new agent tools: burn rates that get away from you too fast. 

As one HackerNews user jabs: “I love when they put a black box in front of the other black boxes so I can get a questionably better black box for slower service and more money!” 

Is collective intelligence the future? 

On X, HA posits that large-scale, monolithic models have had their time in the sun and that solving more complex real-world challenges will require a different beast: collective intelligence. 

Moving forward, Sakana plans to incorporate new models in its agent pool, which could shore up that resilience Sakana is aiming for. But so far, users seem to question whether paying another company to sit between them and frontier models is really worth the spend.

The post Sakana Fugu is more than a router. But it’s not the blueprint for AI sovereignty, either. appeared first on The New Stack.

❌