Normal view

OpenRouter called itself the “Stripe for LLMs” — now Stripe’s swooped in to buy it

Abstract flat-design illustration of thick red, yellow, blue, and green lines intersecting and curving like a subway map, with colors blending into gradients where they cross, depicting model routing.

After weeks of speculation, fintech giant Stripe has confirmed that it’s tabled a bid for AI model gateway platform OpenRouter, a deal designed to help businesses optimize how they route and spend AI tokens.

While terms of the deal have not been disclosed, independent reports peg the acquisition price at a cool $8 billion, making it Stripe’s largest known acquisition to date.

To a casual observer, the deal marks a somewhat odd combination: why would a payments processor want to own technology that decides which AI model answers a given prompt? Well, it all ultimately comes down to “tokenomics” — the emerging discipline of managing the cost, allocation and consumption of AI tokens.

On top of that, OpenRouter has previously said that people should think of it as “like Stripe for LLMs,” owing to the fact that it makes the fragmented AI model market accessible through a single developer-friendly API, much as Stripe did for payments. And that synergy will now culminate in the two companies becoming one.

Token gesture: ‘making good use of scarce compute resources’

Stripe became a $159 billion juggernaut as the developer plumbing behind online payments — the infrastructure that lets internet businesses accept money, run subscriptions, and get paid globally. While its core pitch has always been about making it easy for businesses to accept money, AI has become one of the biggest costs those same businesses have to manage, and managing both sides of that ledger is part of Stripe’s job.

Stripe has been building out AI billing infrastructure long before the OpenRouter deal, previewing LLM token billing and an LLM proxy for routing and metering model calls in 2025. With OpenRouter under its wing, Stripe gains a much more sophisticated routing layer that can choose between hundreds of models and providers based on cost, speed and performance.

In its announcement on Wednesday, Stripe co-founder and CEO Patrick Collison says that “tokens are the central currency for companies building with AI,” adding that the acquisition is ultimately all about the economics of AI.

“Tokens are the central currency for companies building with AI, and it’s clear that the real-world economic potential will depend on making good use of scarce compute resources.”

“The real-world economic potential will depend on making good use of scarce compute resources,” he notes. “Stripe is building the economic infrastructure for AI, and together with OpenRouter we’ll help businesses maximize profitability by routing their requests intelligently and spending their tokens efficiently.”

Open sesame

OpenRouter itself is a relative newcomer to the technology world. Started in early 2023, and co-founded by former OpenSea CTO Alex Atallah, the platform acts as a single front door to the increasingly crowded AI model market. Developers can use one API to access and switch between hundreds of models from dozens of providers, without having to rewrite their applications every time they change models.

Underneath that common interface, OpenRouter handles much of the messy stuff: routing requests between providers, automatically falling back when one goes down, and optimizing for things such as price, latency and model quality. It generally passes through providers’ inference prices without a markup, instead making money through a 5.5% fee on credits purchased through the platform.

That proposition has helped it gain sizeable traction. OpenRouter now says it serves more than 10 million developers and companies across more than 400 models, processing over 10 trillion tokens per day.

OpenRouter
OpenRouter

The company is also fresh off the back of a $113 million funding round, led by Alphabet’s growth fund, with participation from a slew of high-profile backers including the venture arms of Nvidia, Databricks, Snowflake, MongoDB, and ServiceNow — a strategic bet by some of the biggest names in AI and enterprise software.

“AI has become the single largest driver of economic growth in the US, and inference is quickly becoming the largest line item for every company.”

In its own announcement post, penned by founders Alex Atallah, Chris Clark, and Louis Vichy, OpenRouter positions the deal against a bigger shift in where businesses are spending their money: away from simply building AI products and toward the ongoing cost of running them.

“AI has become the single largest driver of economic growth in the US, and inference is quickly becoming the largest line item for every company,” they write.

As for why Stripe, OpenRouter points to a shared developer-first heritage. Stripe’s APIs became something of a benchmark for developer software, while its payments infrastructure gives it experience handling huge volumes of transactions, fraud and abuse — problems OpenRouter increasingly faces as AI usage grows.

The company also suggests that remaining independent was a perfectly viable option, and that very few potential buyers could have persuaded it otherwise.

“There are few companies on earth we would have considered selling to; our mission, our neutrality, and our lead in the market make the story for independence strong,” they write. “We would only join a company if we thought we could do more together, faster, without compromising any of them.”

For customers, OpenRouter’s message is essentially business as usual. Stripe will own the company once the deal closes, but OpenRouter says its brand, product, roadmap and model-neutral approach will remain as is.

“There are few companies on earth we would have considered selling to; our mission, our neutrality, and our lead in the market make the story for independence strong.”

That continuity will likely matter, too, because OpenRouter is far from alone in trying to solve the problem. A slew of companies this year have been investing in their own routing layers, as AI inference costs climb and no single model stays the best or cheapest option for all that long.

The model-routing rush

Cursor Router
Cursor Router

Cursor, the AI coding tool now owned by Elon Musk’s SpaceX, shipped its own Router back in July, claiming savings of 30-50% compared with routing every request through its priciest model.

Ramp, the $44 billion spend-management company, also debuted its very own model router in July, a product that launched on Wednesday at its own dedicated Router.com domain — the same day Stripe announced its deal with OpenRouter. The company says three years of tuning its own AI spend internally cut its bill by 30% — the pitch to new users now promises a bigger number, an average 40% cut.

Meta, for its part, is also reportedly building a model router of its own. The Information reported in July that it’s planning Switchboard — a project out of an internal incubator called AAI Labs, that scores each request for difficulty and routes the easy ones to cheaper models. It’ll stay internal at first, aimed at cutting Meta’s own AI agent bill, but could eventually ship as an external product too.

All this activity speaks to a much broader reckoning over the cost of AI. In June, the Linux Foundation announced the Tokenomics Foundation, backed by the likes of Google, Microsoft, IBM and Salesforce, to develop common standards and benchmarks around how AI tokens are produced, consumed and monetized.

Model routers are one practical answer to the broader underlying problem: spend less by being smarter about which model gets each job. And with OpenRouter now set to become part of Stripe, those economics are moving directly into the payments giant’s wheelhouse.

The post OpenRouter called itself the “Stripe for LLMs” — now Stripe’s swooped in to buy it appeared first on The New Stack.

Arm and Google offer a smarter option to run agentic AI workloads

Warp speed light streaks radiating outward on blue background

As enterprise leaders start deploying agentic workflows, they must establish the infrastructure to build and run them, one capable of fluidly routing a diverse set of workloads across the most efficient compute resources.

This requires the ability to manage heterogeneous infrastructure, utilizing high-performance accelerators for large-scale training and inference, and utilizing CPUs for the critical orchestration layer of agentic AI. As autonomous agents become more prevalent, CPUs are ideally suited for managing agent state, semantic routing, tool selection, and spinning up secure, isolated sandboxes to safely execute untrusted generated code.

The Google Axion advantage

Google Cloud, with its workload-optimized Compute Engine portfolio, which includes general-purpose and specialized offerings, shines in addressing this need.

Google Axion processors within this portfolio comprise a family of custom Arm processors engineered for performance, efficiency, and versatility, with a feature set that supports general-purpose workloads, CPU-based AI workloads, and other specialized tasks requiring Arm-native compatibility and direct hardware access.

Axion is Google’s first custom Arm-based server CPU, introduced in April 2024. It is designed specifically for hyperscale cloud and AI-era data center workloads. 

Axion also leverages more than a decade of Google’s custom silicon innovation. This enables Google to more readily incorporate customer feedback into chip designs and address the more general, though complex, needs of CPUs. 

Matching workload type to the processor

Bhumik Patel, Director of Software Ecosystem Development at Arm, says the key to all of this is to match the workload type as closely as possible to computing capacity. CPU-powered cloud instances are a practical option for certain AI workloads, particularly those with smaller datasets or less complex models. 

“Agentic tasks such as orchestrating, talking to APIs, and memory management are all ones CPUs are good at, so it’s a distributed and concurrent AI workload,” Patel tells The New Stack. Intelligent workload-processing apportionment makes agentic AI more cost-effective and efficient than running all workloads on a single compute type.

This efficiency is quantifiable. The Google Kubernetes Engine Agent Sandbox running on Google Axion N4A provides up to 30% better price performance than the next hyperscale cloud provider, says Google’s Mo Farhat, Axion Group Product Manager. The GKE Sandbox is an open-source Kubernetes-native primitive designed to execute untrusted AI-generated code safely. 

“Agentic tasks such as orchestrating, talking to APIs, and memory management are all ones CPUs are good at, so it’s a distributed and concurrent AI workload.”

Intelligent workload decoupling makes agentic AI significantly more cost-effective. Google Cloud’s fluid computing foundation enables engineering teams to reserve specialized accelerators strictly for heavy reasoning and generative workloads, while leveraging Axion CPUs for high-concurrency orchestration and context management.

Secure execution with the GKE Agent Sandbox

As agents begin to generate and execute dynamic code autonomously, security is non-negotiable. Running AI-generated code directly in a standard cluster poses severe security risks, as untrusted code could potentially access other apps or the underlying cluster node.

The Google Kubernetes Engine (GKE) Agent Sandbox resolves this by providing an isolated environment for safely executing untrusted code. Running on Axion-powered N4A instances, the sandbox provides up to 30% better price performance than comparable workloads on other hyperscalers.

The vertical stack isolates sensitive tasks at the kernel level with sub-second latency.

The vertical stack isolates sensitive tasks at the kernel level with sub-second latency.  GKE Agent Sandbox natively supports gVisor (an open-source application kernel developed by Google that acts as a secure sandbox for containers) and default-deny Kubernetes network policy. Agent Sandbox provides pluggable interfaces for open-source sandboxes, such as Kata Containers, enabling users to customize their kernel isolation. 

Powered by gVisor technologies with software support from Arm’s architecture, the sandboxes intercept and validate system calls before they reach the host kernel. These isolated execution environments enable deployment of autonomous systems at scale without sacrificing performance or operational agility.

To manage resources efficiently when agents sit idle, GKE Pod snapshots allow users to save and restore the exact process state of sandboxed environments. This functionality provides four major architectural benefits:

  • Fast startup: Reduces sandbox startup time by restoring from a pre-warmed snapshot rather than initializing from scratch.
  • Long-running agents: Pauses sandboxes that take a long time to run and resumes them later—or moves them across nodes—without losing progress.
  • Stateful workloads: Persist an agent’s context, such as conversation history or intermediate calculations.
  • Reproducibility: Captures a specific state to use as a baseline for spinning up multiple new sandboxes.

Getting started

As token generation, autonomous workflows, and continuous agent interactions grow exponentially, relying exclusively on accelerator-backed stacks for every task will become financially and architecturally unsustainable.

The combination of CPU and accelerator execution accounts for bursts in agent activity and unpredictable demand spikes by eliminating the inference tax. Google Cloud’s full-stack advantage enables organizations to deploy the right machine for the job. 

By using Google Axion and GKE Agent Sandbox, builders can optimize total cost of ownership and security while maintaining the performance required for AI agents.

Learn more about Google Axion.

The post Arm and Google offer a smarter option to run agentic AI workloads appeared first on The New Stack.

Meta and the rise of the accidental cloud

Close-up of server rack hard drive bays with yellow locking handles

In the same quarter, a $1.7 trillion social network and a shoe company both became cloud providers. Nobody’s operating model was designed for this.

On July 1, Bloomberg reported that Meta was building a cloud business to sell its excess AI capacity, signaling that one of the largest GPU fleets on Earth is about to get a price list.

As the report outlines, the company is still weighing two business models that are part of its Meta Compute initiative: it could offer hosted access to AI models running on Meta infrastructure (similar to AWS Bedrock), or directly rent raw compute capacity, à la the CoreWeave model.

Two weeks before that report, Allbirds had finished becoming Smartbird. The erstwhile sneaker company — which hit its peak in 2022 in selling 3 million pairs — sold its footwear brand for $39 million, lined up a convertible note facility that it later expanded to $100 million, hired an ex-AWS executive as CEO, and set out to sell GPU-as-a-Service. 

When it first announced the pivot in April, the stock spiked nearly 600% in a day, adding more than $100 million in market value at its peak — all before the company had racked a single GPU.

One of these is a serious supplier, and the other is a public shell chasing an AI multiple, but I don’t think the distinction matters much. Compute supply is fragmenting faster than any enterprise can absorb it, and overbuild always finds a buyer.

Overbuild becomes inventory

Meta isn’t selling compute because its leaders woke up wanting to fight AWS. It’s selling compute because it provisioned for its own peak, and there’s a gap between what it built and what it uses. That’s what an accidental cloud is: Infrastructure that was never meant to be a product, monetized because the alternative is depreciation.

That’s what an accidental cloud is: Infrastructure that was never meant to be a product, monetized because the alternative is depreciation.

Meta won’t be the last. Everyone who bought more GPUs than they needed in the last three years is facing the same math. Some will eat the write-down. The rest will sell. Stack that on top of the neoclouds (CoreWeave, Nebius, Lambda), the sovereign clouds, and now the Smartbirds, and the list of places you can buy serious compute has gone from a handful to many in about two years. 

The market read the Meta news as bad for neoclouds. CoreWeave and Nebius both dropped double digits on the day, and I get why: Nebius has a $27 billion contract with Meta and CoreWeave a $21 billion deal. Both had Meta as a customer and later witnessed it become a competitor in the compute business. But I’d pull a different lesson from it. It’s not that one supplier wins and another loses. It’s that supplier positions are now unstable everywhere. 

If the suppliers can’t predict their position two years out, betting your operating model on any one of them is a risk you’re not pricing.

More suppliers should mean leverage. Mostly it means sprawl.

On paper, a fragmenting supply side is great for buyers: price competition, more choice, more leverage. Cheaper compute is coming, and for AI workloads, it’s coming fast.

Most teams I talk to can’t capture any of it. Every new supplier shows up with its own console, its own billing format, its own identity model, and its own hole in your governance coverage. Signing a supplier is cheap. Operating one is not. The security review, the tagging standards, the budget enforcement, the offboarding plan — all of it gets rebuilt per provider. Choice without governance isn’t leverage. It’s sprawl, and sprawl costs more than the discount that created it.

Choice without governance isn’t leverage. It’s sprawl, and sprawl costs more than the discount that created it.

This bites hardest with GPU capacity, because that’s where the fragmentation is happening and where the money is. Workloads end up pinned to whichever supplier had chips available the day the contract was signed, and they stay there. Not because moving is impossible, but because nothing above the suppliers makes moving routine, and the team that signed the contract usually isn’t the team on the hook for utilization. Those incentives don’t fix themselves.

The Smartbird end of the market makes this non-optional. Capacity from a vendor with no enterprise track record is only usable if you can exit it in a day. That’s something you design for up front, not something you negotiate into a contract.

The durable position is above the suppliers

Every argument about picking the right cloud assumes the list of clouds is stable. It isn’t, and it’s about to get less stable. The position that survives supplier churn is the layer above them: a single control plane where every provider — hyperscaler, neocloud, or accidental cloud — is just a target you provision to, under the same policies, approvals, cost visibility, and exit path.

This is the VMware argument, extended forward. Broadcom taught the industry what single-supplier dependence costs when the supplier’s incentives change. Meta just taught the follow-up lesson. The future holds more clouds, not fewer, arriving faster and from stranger directions than anyone planned for.

The enterprises that win the price war won’t be the ones that picked the right supplier. They’ll be the ones for whom the supplier stopped mattering.

The practical version of this is an abstraction layer that treats every supplier as a provisioning target rather than a separate operating model. When a new supplier appears — Meta Compute, or a neocloud that didn’t exist last quarter — it plugs into the governance the team already defined, rather than becoming a new operating model to build from scratch. The policies get written once; the supplier list underneath can churn.

Meta selling compute and a sneaker company selling GPUs are the same headline: supply is no longer scarce; it’s fragmented. The enterprises that win the price war won’t be the ones that picked the right supplier. They’ll be the ones for whom the supplier stopped mattering.

The post Meta and the rise of the accidental cloud appeared first on The New Stack.

Public cloud vs. on-prem: Summit on where each workload belongs

On this episode of The New Stack Makers, Summit’s Byron Dill argues that many enterprises have become overly reliant on public cloud infrastructure, using it for workloads that may be better suited to private environments.

We’re more than 20 years past the launch of AWS, the starter gun for the shift of compute and storage from on-prem racks to the cloud. 

The rapid growth of AWS and competing services like Azure and Google Cloud underscores how many companies have made the jump from controlling their own infrastructure to renting capacity from hyperscale public clouds.

For the major providers, the public cloud has proved an incredible business. Amazon’s cloud service generated nearly 60% of its first-quarter operating profit, for example. For cloud customers, however, the tides may be turning.

Think back to the early days of the public cloud. Azure and AWS scrapped for market share, offering price cuts to entice workloads to their centralized silicon. The situation has evolved over the ensuing decades. Today, cloud costs are material and rising, prompting some companies to question whether being cloud-first is the best path forward.

Cloud bills are expanding due to increased usage of hyperscaler infrastructure, yes, but also because many customers today use the cloud for everything, rather than for what it is best suited for.

n the latest episode of The New Stack podcast, Byron Dill, Director of Solutions Engineering at Summit, tells us that shared compute and storage have their place in the modern IT mix, but that many companies would do well to segment their workloads and move some of that work back on-prem. (Think lower costs and simpler management of high-risk data.)

The argument echoes what we’ve seen recently in the AI realm. Many companies quickly adopted AI technology, only to be surprised later by the bills they incurred. The public cloud is a similar frog-boiler, albeit on a slightly longer timeframe.

In both cases — AI and the public cloud — companies have learned that a product once pitched as a way to reduce spend can evolve into the opposite without careful management. Summit, which offers managed private clouds to enterprise customers, thinks that some corporate workloads should be removed from the cloud and moved in-house.

What will that cost? How long does it take to move? And which industries are most primed to benefit from their own private cloud? We get into it all in this episode.

The post Public cloud vs. on-prem: Summit on where each workload belongs appeared first on The New Stack.

Kubernetes teams trust automation to ship code but not to touch CPU, and AI is raising the stakes

Kubernetes teams automate deployments without thinking about it. CI/CD pipelines fire dozens of times a day, autoscaling adjusts replicas in the background, rollback is muscle memory. But there is one category of automation where that confidence vanishes: letting a system change CPU and memory requests on a running workload without a human reviewing it first. 

And as AI inference lands on Kubernetes at scale, that hesitation is becoming hard to ignore, and increasingly expensive.

Why teams trust automation for change but not for constraint

We surveyed 321 Kubernetes practitioners at enterprise organizations earlier this year. The headline finding is one most practitioners will recognize immediately: 82% report high or complete trust in automated delivery controls. But 71% still require human review before applying resource optimization recommendations. Only 27% allow CPU and memory changes to be auto-applied, even within guardrails.

“Deploying code feels additive… rightsizing feels subtractive because you are removing safety margin from a running service, and the failure mode is fundamentally different.”

Those numbers describe a specific asymmetry. The same engineers who deploy to production dozens of times a day without hesitation slow down the moment automation wants to adjust resource allocation. And the survey data make it clear why. Deploying code feels additive. You are shipping new value, the rollback path is well understood, and if something breaks you usually see it right away. Meanwhile, rightsizing feels subtractive because you are removing safety margin from a running service, and the failure mode is fundamentally different.

As one practitioner in the survey put it: “Automated right-sizing carries a unique risk because it directly impacts the underlying stability of the application runtime. Unlike a code deployment that follows a tested path, resource changes alter the invisible contract between the workload and the scheduler.”

When you change resource requests, you change how Kubernetes schedules, prioritizes, and allocates resources. Those effects are not visible the way a code change is. You can’t trace them through a deployment pipeline. And you might not discover that something went wrong until two weeks later, when a traffic spike hits a threshold that didn’t exist at the old values. By that point, three other things have changed too, and proving causation is nearly impossible. The people responsible for those workloads are the same people who get paged at 2 a.m., and they know this.

Why AI workloads raise the stakes

That trust gap existed before inference workloads showed up. What’s changed is the cost of not closing it.

For a long time, teams could absorb the cost of manual oversight. They knew their workloads, had intuition for where the safe boundaries were, and the inefficiency of over-provisioning was a price worth paying for stability. GPU-accelerated inference workloads change that math. GPU compute is significantly more expensive per hour than CPU. The cost of over-provisioning is no longer a rounding error you can absorb quietly. And the workload behavior is less familiar, as inference jobs are bursty in ways teams haven’t built intuition for, traffic patterns shift as models are updated and usage changes, and the resource dimensions involved differ from what teams have spent years learning to tune.

That unfamiliarity compounds with scale. Rightsizing isn’t a one-lever problem the way horizontal scaling is. It involves, at minimum, CPU and memory requests and potentially limits for both, with four dimensions per workload, multiplied across hundreds or thousands of workloads per cluster. The survey data indicates that manual optimization breaks down at around 250 changes a day. Inference workloads push teams past that threshold faster than anything they’ve managed before, because the resource decisions are more frequent and the cost of getting them wrong is higher.

The economic case for automated rightsizing has never been stronger. The organization’s willingness to delegate hasn’t caught up because teams are being asked to trust automation with workloads they don’t yet have a track record with.

What the survey says about closing the gap

When we asked practitioners what would actually increase their trust in optimization automation, 48% said visibility and transparency into how decisions are made, 25% wanted proven guardrails, and 23% needed instant rollback.

Nobody asked for full manual control and very few asked for blind autonomy. What they described is automation that earns trust in stages, and that’s consistent with how the teams furthest along in their automation journey actually got there. They didn’t start with production. They started with a single namespace in a dev environment, observed the system’s behavior, compared recommendations with outcomes, and gradually expanded the scope. Different environments remained at different levels of automation maturity simultaneously, and that was intentional. Production carried more scrutiny than dev.

CI/CD followed the same curve, and the timeline is easy to forget. Most organizations took years to get from running their first automated pipeline to trusting it with production deploys without manual approval on every commit. Kubernetes resource automation is earlier in that same process, and AI workloads are extending the timeline because teams are building trust from scratch with a workload category that doesn’t yet have a track record.

Why automation design matters as much as capability

Some automation architectures deliver meaningful value only with full delegation. The system needs complete control to function the way it was designed to. That’s a form of forced autonomy, and it creates an adoption problem because it asks for exactly the level of trust that most organizations haven’t built yet. Force generally doesn’t work. Teams that feel pushed into a level of delegation they aren’t comfortable with tend to pull back entirely after the first incident.

The alternative is what I’d describe as adaptive autonomy: designing the system to work at every stage of the trust curve. A team still evaluating gets useful recommendations in read-only mode. A team ready to act but wanting boundaries can run guardrailed execution within limits they define. As confidence grows, the system handles more decisions autonomously while humans manage exceptions. And for environments where the track record supports it, closed-loop optimization runs in the background and becomes boring, which is the goal. Each stage is a legitimate operating mode, not a stepping stone you have to rush through.

That design distinction matters more with AI workloads than it ever did with traditional services, precisely because the trust-building process is starting from zero on workloads where the cost of getting it wrong is highest.

“Trust takes a long time to build and a single production incident to undermine.”

The other piece that makes this sustainable is rollout safety. Trust takes a long time to build and a single production incident to undermine. Start with the workloads showing the most headroom between requests and actual usage. Make changes incrementally, small enough that a bad outcome stays contained. Rollback needs to be fast and tied to the health signals the team already monitors. And start with opt-in, not opt-out. Let the teams willing to go first build a track record that others can look at.

The broader pattern

The 71% figure is sometimes read as resistance to automation. I think it’s a more accurate picture of how operational trust actually forms: conditional, earned over time, and moving at different speeds depending on what’s at stake. AI workloads are raising those stakes significantly, which means the path to trusted automation matters more now than it did when the cost of caution was just some unused CPU.

“Most of what gets written about Kubernetes optimization focuses on tooling capability, and the tooling is capable. The harder problem is the human one.”

Most of what gets written about Kubernetes optimization focuses on tooling capability, and the tooling is capable. The harder problem is the human one. If your team is managing AI inference workloads on Kubernetes and your optimization tooling is sitting in read-only mode, the question worth asking isn’t whether to trust the system. It’s whether the system is designed to let you build that trust gradually, starting where the stakes are low and expanding as the evidence supports it, on workloads where getting it wrong costs more than it ever has before.

The post Kubernetes teams trust automation to ship code but not to touch CPU, and AI is raising the stakes appeared first on The New Stack.

How AI is solving the memory crunch it created

Close-up macro photograph of computer RAM memory modules and a circuit board lit in red and green, showing gold contact pins and electronic components.

Memory has replaced compute as a primary constraint for modern tech teams. A perfect storm of hardware architecture limitations, semiconductor supply chain uncertainty, and changing software licensing models has left enterprises confronting increasingly memory-constrained environments. All while high-bandwidth AI workloads overload the production chain’s ability to provide sufficient memory — and when your AI token bill is starting to cost more than your salary bill.

Over the past year, the cost of high-bandwidth memory (HBM) and dynamic random access memory (DRAM) has increased by an unprecedented 170%, with some virtualization subscriptions more than doubling in price.

All this adds up to a demand for enterprises to shift from the previous buy-all-you-can mindset to a data-driven optimization strategy.

Fortunately, AI isn’t only part of the problem. When AI is applied to memory economics in modern virtualization, it becomes a vital part of the solution. 

Bharath Ram, director of product management at Hewlett-Packard Enterprise (HPE), explains it to The New Stack this way: “There’s a component shortage. Today the prices have increased. So customers are looking at ways to save and optimize the existing footprint, so that they can run their workloads on whatever and not have to procure anything new.”

By switching focus from gobbling up every bit of memory your organization can grab to optimizing your workloads and their placement across multi-cloud and hybrid-cloud environments, enterprises cannot only speed up modernization but also shorten decision-making cycle time by up to 80%, all while cutting costs by up to 50%. Read on for how to transition your enterprise from guesswork to smarter IT.

Tech has to confront its waste problem

Most enterprises operate with significant over-provisioning driven by limited visibility and risk avoidance.

These same companies are standing up legacy applications that become less efficient over time, including a significant number of zombie services running without any use. 

On top of this, most AI workloads rely on advanced memory technologies, which has led chip manufacturers to shift production priorities from DDR4 to DDR5 RAM, further reducing DDR4 supply. This impacts the whole industry, with even a personal computer costing 15% to 30% more than last year.

Add to this volatile DRAM pricing and higher core densities, and it’s clear that even non-technical leadership is worried about memory efficiency. Tech giants Microsoft, Google, Amazon, and Meta are buying up as many AI chips as they can, which is triggering even more shortages and price pressure across the supply chain. And thus more enterprises are hoarding more infrastructure and memory. 

And this overbuying isn’t limited to memory. Companies are now also buying infrastructure like servers even before they need them, too, Ram reflects, “because the cost is so exorbitant, the quote that you might have today might not be the same price that you’re quoted for the same infrastructure tomorrow. That’s how we’re seeing the market right now. It’s very volatile.” 

But it might all be ok. HPE estimates that between 20% and 40% of infrastructure is overprovisioned today. Which is an opportunity for efficiency — not only in these limited resources but also in faster, more secure workloads. 

Enterprises are more capable than ever to optimize the use of what they’ve got today, especially before they go searching for more RAM that will cost significantly more.

It all starts with understanding

So much of this waste persists because enterprise infrastructure is obscured — no one really knows what does what with which data, or which services rely on it. 

The same thing that holds companies back from doing anything more than lift-and-shift to the cloud is usually what keeps them from unlocking memory efficiency. There’s simply too little visibility across most enterprises’ complex, hybrid and multi-cloud distributed systems. Which has left organizations guessing and then rounding way up for over a decade now.

“It’s a combination of over-provisioning and not understanding underlying infrastructure. Because many of them are doing public cloud-based provisioning and self-service, where you don’t know what the underlying infrastructure is and you have admins leveraging whatever there is in terms of their service capabilities,” Ram explains. “One piece is memory shortage, and the other is understanding what’s been deployed and rightsizing it.”

To break these over-provisioning bad habits, any change has to be grounded in reality. The first step is to gather and analyze real usage data, using a tool like HPE CloudPhysics to establish a factual baseline that separates real cost drivers from those years of assumptions. 

This allows enterprises to:

  • Understand their virtualization footprint and licensing exposure.
  • See their workload initialization and efficiency.
  • Identify true cost drivers before taking action.

You cannot right-size until you have real-time monitoring of how many hosts have how many VMs, and which are on and off.

Predictive, not reactive provisioning

Once an enterprise has a single source of truth for its complex distributed systems, it can explore what to deploy, where, when, and how.

“An application like SAP HANA is highly memory-intensive and highly latency-intensive. It’s not like this algorithm is optimized to pivot between hot and cold memory tiering” for cost reduction, Ram explains, without risking the application performance, akin to how, when older PCs had limited amounts of memory and, once that ran out, the computer would swap the program from running in memory to disk, slowing way down. 

Part of the modern solution, Ram argues, is that companies “can over-provision with what they already have. They don’t have to buy any new memory,” because of better shared resources available to all the virtual machines managed by a single host. 

“For example, a host with 64GB of physical memory may have more memory allocated across VMs than physically available,” he explains. “In practice, not all VMs consume their full allocation simultaneously, allowing unused capacity to be dynamically reassigned where needed.”

Memory ballooning, which, Ram says, is nothing new, but something desperately needed in the market right now. Version 9.0 of Morpheus, due out this summer, will feature a more modern sort of memory oversubscription, which, HPE explains, allows administrators to oversubscribe physical memory across VMs on a host, enabling higher VM density and more efficient use. This is particularly useful for testing and development environments, virtual desktop infrastructure, and workloads with variable memory demands.

Shift to architectural efficiency 

Eventually, once you’ve optimized and rightsized every memory allocation, it’s time to shift your workloads to a new platform to improve hardware efficiency. 

“The final step is increasing workload density per server, especially as per-core software licensing becomes more expensive,” Ram explains. He says this is best achieved using a virtualization solution with an open-source hypervisor, which can improve utilization now and help organizations shift toward per-socket licensing models. For suitable applications, that modernization may also include moving to containerized deployment, while in-memory deduplication reduces redundant data structures in RAM, improving memory efficiency.

“Not everybody can keep running on existing hardware forever. At some point, some organizations will need to move to a new platform to improve hardware efficiency. But higher workload density still brings added benefits,” he continues, especially at a time when even the biggest tech companies are overbuying infrastructure, driving up costs and tightening capacity.

Per-socket licensing saves more

And if you do go for a hardware refresh with HPE’s Morpheus Software, then you can unlock a different kind of subscription model, which charges per socket or CPU licensing, where multiple cores can share one socket. Some early results indicate that this can deliver up to 90% in savings.

In the end, it all starts with that baseline. Take the free cloud visibility assessment to see where your organization stands.

The post How AI is solving the memory crunch it created appeared first on The New Stack.

The fix for soaring AI cloud bills exists — so why won’t we trust it?

To hear Yasmin Rajabi, chief operating officer at CloudBolt, tell it, there’s an imbalance in how we view automation. We’re happy to automate decisions that result in more productivity and processes — but what about when it comes to turning the dial to the left? For some reason, there’s hesitation there.

“Trust is super-high when it comes to traditional automation, but there’s still a lot of caution when it comes to right-sizing,” Rajabi tells The New Stack. “The same engineers who are deploying multiple times a day through CI/CD aren’t questioning [automation] anymore, but when it comes to delegating right-sizing to the machine, the bar to earn that trust is much higher.”

The data reveal why this imbalance might exist: When faced with the pressure to remain always-on, a higher cloud bill from over-provisioning seems worth the cost. But now that GPU-heavy AI workloads have sent cloud bills soaring, right-sizing automated processes has become a priority for 89% of organizations, according to the March 2026 CloudBolt Research Report.

And yet, 71% of Kubernetes engineers respond that they still require human review for resource optimization, with only 27% allowing CPU and memory changes to be auto-applied. So while the data shows it’s a priority, that motivation hasn’t shown up in the workflows.

“When it comes to delegating right-sizing to the machine, the bar to earn that trust is much higher.”

The New Stack will sit down with Rajabi and Reid Vandewiele, product lead at StormForge, at 9 a.m. Pacific/5 p.m. BST on Wednesday, June 24 to discuss the urgency of this right-sizing gap — especially when it comes to Kubernetes workloads for AI.

Join us live to not only learn how to measure your organization’s automation maturity, but to develop this trust over time, with strategic CPU throttling, out-of-memory (OOM) behavior, and, of course, rollback patterns. 

Register to join this conversation

Right-sizing is a multi-dimensional problem, Rajabi explains, spanning increasingly complex workloads in increasingly complex environments, so that when something goes wrong, it feels almost impossible to reverse. For this automation to work, this trust has to be built not just across teams but scaled across your organization.

“It takes a long time to build up trust in an automation solution, and it’s very fast to eliminate or significantly undermine that trust,” Rajabi warns. “It takes one production incident to take an application team from being willing to entertain automated resourcing to absolutely not, ‘not on my application, we’re special’.”

Join us June 24 to learn how to gain insight into how much your AI workloads actually cost, and to adopt a plan that takes the guesswork out of your provisioning.

The post The fix for soaring AI cloud bills exists — so why won’t we trust it? appeared first on The New Stack.

❌