A recent GigaOm CxO Decision Brief explores how AI retrieval architectures are evolving beyond flat vector databases as organizations combine semantic search, ranking, personalization, and machine learning inference in production systems.
Vector search changed the AI infrastructure landscape by making semantic retrieval practical at scale. By converting text, images, and user behavior into embeddings, organizations could move beyond exact keyword matching and retrieve information based on meaning. But production AI systems rarely stop at vector similarity.
A real-world query often requires multiple signals to be evaluated simultaneously. Semantic relevance may be one factor, but so are structured attributes, business rules, personalization signals, freshness, access controls, recommendation logic, and machine-learned ranking models. As organizations move from AI experimentation to production-scale applications, the challenge is no longer simply finding similar items. It is in combining all of the signals that matter while maintaining low latency and operational simplicity. This is where tensors are attracting increasing attention.
While vectors represent information as a single dimension of numerical values, tensors provide a more general framework for representing and operating on complex, multi-dimensional data structures. They offer more control in how relevance is computed, allowing dense embeddings, sparse features, metadata, and model outputs to be evaluated together within a unified retrieval and ranking process. For organizations building large-scale retrieval systems, this raises an important architectural question: is a flat vector store sufficient, or does the next generation of AI applications require something more expressive?
“Tensors provide a more general framework for representing and operating on complex, multi-dimensional data structures.”
A new GigaOm CxO Decision Brief, “The Tensor Advantage in AI Search,” explores this question in depth.
Among the findings:
Production AI systems increasingly depend on combining semantic, lexical, behavioral, and business signals rather than relying on vector similarity alone.
Architectural fragmentation between vector databases, search engines, rerankers, and feature stores introduces latency, operational complexity, and synchronization challenges that become more significant as workloads scale.
Emerging retrieval models, including multi-vector and late-interaction approaches, place new demands on infrastructure that were not anticipated when first-generation vector databases were designed.
Tensor-native architectures provide an alternative approach by treating multidimensional data structures as first-class citizens rather than forcing them into simpler vector abstractions.
The paper also examines the infrastructure, operational, and organizational implications of these architectural choices, including benchmark data, deployment considerations, and the trade-offs engineering leaders should evaluate when planning future AI retrieval systems.
“Retrieval is evolving from a nearest-neighbor problem into a ranking and decision-making problem.”
As AI applications become more sophisticated, retrieval is evolving from a nearest-neighbor problem into a ranking and decision-making problem. Understanding the role tensors play in that transition may be one of the most important architectural discussions facing engineering leaders today.
Every enterprise software vendor is currently selling some version of the same thing: AI agents grounded in enterprise context and governed by a central control plane. SAP, ServiceNow, Salesforce — they all have one.
At its ONE conference in Amsterdam in June, OutSystems unveiled its version, and its CEO, Woodson Martin, agrees that they all look quite similar on the surface, but unsurprisingly, he also believes that OutSystems has a very different approach.
Martin tells The New Stack that, in his view, “It would be very easy to just look at the market today and say all enterprise software players are offering exactly the same thing.” The reason for that, he says, is that enterprise agent orchestration, “is sort of greenfield today. Everybody’s aiming for it. Everybody’s got a great story about why they’ll be a leader or a player.”
For OutSystems, the story is neutrality. While SAP and Salesforce pitch agent orchestration from within their ecosystems, where they are also the systems of record, the 25-year-old former low-code company, which now describes itself as an agentic systems platform, wants to be the layer that coordinates across all of them without owning the underlying data.
OutSystem’s agent platform. Credit: The New Stack
The advantage of not being a system of record
Martin says the company has been playing some version of this for a long time. “We’re glue between commercial off-the-shelf solutions,” he says. “We’re the thing that makes the enterprise their own enterprise, as opposed to an SAP enterprise or a Salesforce enterprise.”
One asset management customer, he says, has used OutSystems as the orchestration engine across about 80 systems for fund onboarding for the past six or seven years. That wasn’t for any agentic systems yet, of course, but OutSystems’ role in this isn’t all that different. “We’re already playing that orchestrator role,” Martin says. “In other cases, we don’t have that position yet in the account, and we’ll have to fight for it.”
Tiago Azevedo, OutSystems’ CIO, makes the same argument. “We are agnostic to all of those things.” The OutSystems platform doesn’t create most of the data it touches, he notes. Instead, its focus was always on integrating existing systems. “Our happy place is when we bring several of those systems all together into a form that makes sense for a process,” he says.
Credit: The New Stack.
Open to Claude, Codex, and Kiro
At the ONE conference, the company launched the OutSystems Agent Experience, a platform layer that exposes Model Context Protocol (MCP) and Agent2Agent (A2A) services. Developers can now build, publish, and extend OutSystems applications with third-party coding tools like Claude Code, Codex, Cursor, and Kiro, AWS’s spec-centric IDE.
The first of those services is now live on the OutSystems Developer Cloud (ODC), the cloud-native, current-generation platform. Support for OutSystems 11, the older self-managed platform where a large part of the installed base still runs, launched in early access
“Existing customers that live on O11, they’re like: what I would love to do was to be able to use Claude or Codex or whatever to evolve my applications in O11,” Azevedo says. “So we made that possible. “[…] We made a lot of people happy.”
And indeed, when this was announced in the keynote, it drew more applause than some of the other large product announcements.
He sees no alternative to opening up the platform, given how freely developers now move between coding tools. “I strongly believe in open systems,” he says. “You close those environments, you’re gone.”
Other launches at the conference include the Agentic Enterprise Orchestration service and the next-generation OutSystems Agent Workbench, which is now generally available and adds agent evaluations, guardrails, semantic search, and Amazon Bedrock support.
There is also a preview launch of a new modernization service, built on AWS Transform and Kiro, for migrating COBOL and Lotus Notes systems onto the platform, as well as a pre-packaged agentic solution for loan origination, the first of a family of packaged agentic industry solutions, which arrives later this year.
The new bane of IT departments: shadow AI
Shadow IT has returned as shadow AI, Azevedo says. And while he has always managed to stay ahead of internal technology demands with previous platform shifts, that’s getting harder now. “With AI it’s impossible,” he says. “It’s literally impossible. It’s not humanly possible.”
The way he describes it, a central team can build maybe 10 large agentic workflows that solve company-scale problems. “Those are what we call the big bets,” he says, and these bets exhaust the team’s capacity. Everything else means letting the rest of the organization build its own agents, and every one of those requests immediately raises questions of who gets to touch which company data and through which MCP servers. Demand from every department, he says, is growing almost exponentially.
The token bill
Unsurprisingly, this is now also coupled with the question of how much all these tokens cost.
“I also have my CFO saying, what about the token usage? And what about the budget? And who’s gonna pay for that?” he says. “If you look at, let’s say, 500 euros or dollars a month, times 12, times, let’s say, 1,200 or 1,500 people, these are millions in a year.”
For now, at OutSystems, he rations token budgets almost by hand, protecting projects that could scale and trimming those that serve an audience of one.
“It’s the most expensive software — and I managed big contracts,” says Azevedo. “The most expensive I’ve ever had in my hands.”
“It turns out my number one consumer of tokens on Anthropic in the month of April was a business value consultant in Australia … why is he burning $7,500 in tokens every week? That’s not in the budget.”
Martin tells a similar story from the CEO perspective. “It turns out my number one consumer of tokens on Anthropic in the month of April was a business value consultant in Australia,” he says. “And we’re like, why is he burning $7,500 in tokens every week? That’s not in the budget.”
Martin traces the shift to the latest generation of reasoning-heavy models, which arrived around January and February, “and we’re getting these token bills in March and April that are starting to scare everyone.”
For the overall OutSystems platform, the answer to this is model flexibility. Customers can bring their own models, swap them without touching the agent logic, and route requests through Amazon Bedrock to whatever is the cheapest option to do the job effectively. Some customers built model routers on OutSystems early on, Martin says, sending complex asks to expensive models and simple ones to cheaper ones.
He also argues that OutSystems’ Enterprise Context Graph means that reasoning over an enterprise context graph is less token-intensive than reasoning over an application’s raw codebase.
“I think organizations are going to develop more discipline around this as we have for other elements of material spend in any enterprise,” Martin says. “This is now becoming material for almost everyone.”
OutSystems’ advantage may indeed be that it isn’t a system of record but ties into all of them. As those systems of record open up with new MCP tools, the former low-code platform may just be in the right position — and with the right customer base — to help its customers tie all of these together, even as those platforms launch their own agent builders and orchestration platforms. Sometimes you may want your agent to be close to the data, but over time, nobody wants to manage half a dozen agent orchestration platforms either.
Beneath the chatbots and copilots, there’s a quiet revolution happening in the data services space. From pure-play database vendors to data integration wranglers and onward to the cloud hyperscalers, the focus has shifted.
Now in the spotlight is the question of how to automate data governance for agentic AI workloads, and for good reason: Traditional manual data stewardship doesn’t scale in a world where agents are becoming increasingly autonomous (and powerful).
Aiming to cut a swath in this marketplace is data control plane company lakeFS. The organization announced its lakeFS for Agentic AI service on Wednesday, and it appears to be designed to bring governed, reproducible data access to autonomous and headless agentic workloads (those that execute decisions below the user interface level) that run at enterprise scale.
The manual model breaks
Einat Orr, CEO and co-founder of lakeFS, tells The New Stack that manual data stewardship was built for human-paced, human-reviewed workflows, i.e., someone looking at a change before it is committed.
“When dozens or hundreds of agents are making changes simultaneously, faster than any person can review, the manual model breaks,” Orr says. “This is because with a human analyst, a bad write to production is usually one mistake, caught by another human before it spreads far. An agent is different — it acts automatically, in parallel, at machine speed, and it doesn’t pause to second-guess itself. And because so much agent activity is unsupervised, you often find out after the damage is done.”
She explains that attempts to identify and roll back incorrect or corrupted production data across a wide set of data modalities, such as images, documents, metadata, and structured data, are almost impossible to pull off. Impossible, that is, unless the team has the data infrastructure in place to isolate and track such changes automatically.
While some of the more disastrous outcomes stay inside an organizaton’s perimeter (or are swept beneath the communications radar), Orr explains that real world consequences of bad agentic data writes are manifold.
“Insurance claims get inappropriately denied or approved, sensor data from machines gets misinterpreted, an incorrect medical diagnosis is made, or customer service bots provide incorrect answers to customers,” Orr says. “The cost of an individual action may be manageable, but agents performing these actions hundreds or thousands of times can have an exponentially larger impact.”
“As agents are let loose on enterprise data at a massive scale, any agent that reads or writes to production data without isolation or a reproducible trail is a liability, no matter how good the model is,” —Einat Orr, lakeFS CEO.
Bad agents acting in the real world
Examples of this happening include the July 2025 Replit AI coding agent incident, which deleted a live production database during an explicit code freeze, wiping records for more than 1,200 executives and around 1,200 companies. To tidy up its handiwork, the agent then fabricated thousands of fake records and initially claimed the deletion couldn’t be rolled back.
Also in July 2025, Google’s Gemini CLI agent misread a single failed command, acted on a version of the file system that existed only in its own interpretation of the scenario, and permanently destroyed a user’s project files. The Gemini agent is widely reported to have said of its actions: “I have failed you completely and catastrophically. My review of the commands confirms my gross incompetence.”
“The pattern in both is the same: An autonomous agent took a destructive action that no one authorized, and the lack of isolation and a reliable rollback path turned a single mistake into permanent loss,” Orr says.
A doctor of mathematics with a track record in hardcore software engineering, the bottom line for Orr is clear: “As agents are let loose on enterprise data at a massive scale, any agent that reads or writes to production data without isolation or a reproducible trail is a liability, no matter how good the model is,” she said.
“…any agent that reads or writes to production data without isolation or a reproducible trail is a liability…”
Gartner expects 40 percent of enterprise applications to have task-specific agents embedded by the end of 2026, up from less than 5 percent a year earlier. IDC projects that agent use at the largest enterprises will grow tenfold by 2027, with the API and data calls those agents make growing a thousandfold. That’s the scale production data has to withstand, and it’s what lakeFS is built to govern.
Agents sent to play in an isolated data sandbox
To address these issues, lakeFS for Agentic AI gives every agent its own isolated data sandbox with a “zero-copy” branch of relevant data, so the agent can access the dataset it needs via references, snapshots, or copy-on-write techniques.
This means any changes the agent wishes to make must be validated and merged in accordance with the policy guidelines defined by the system architecture. In turn, this produces a unified audit trail across every agent action.
When running, lakeFS for Agentic AI is powered by its data version control architecture, which provides zero-copy data sandboxing. This enables isolation so that agent mistakes are automatically isolated and never corrupt production data. Every agent run is tied to an exact, immutable version of the data. Past actions can be recreated, debugged, audited, or extended using the same inputs.
Production data is gated by policy. Merges into production happen only after pre-merge validations pass. Every change can carry an agent identity, a run ID, and an execution context. The result is a unified audit trail instead of evidence scattered across orchestrators, model providers, and cloud logs.
Agents confined by branch-scoped credentials
Where agents are permitted to read and write through standard file operations. lakeFS provides file-level data access with branch-scoped credentials. These can be described as strictly cryptographically bounded, ephemeral access tokens that confine an agent to a specific branch of data or code, so that the agent operates only within its own workspace. This whole mechanism keeps each agent’s working set narrow and avoids context bloat.
“With lakeFS Mount, a branch, or even a subset of a branch, can be mounted as a local directory inside the sandbox or virtual machine where the agent is running,” Orr confirms. “From the agent’s perspective, it’s just reading and writing to files and folders.” She further clarifies and notes that no LLM tokens are spent learning the lakeFS API. The agent works with a familiar filesystem interface, and lakeFS handles the versioning underneath.
Developers also have a couple of options for injecting custom validation logic. CEO Orr explains that software engineers can use webhooks or Lua scripts, both of which allow users to define behavior and rules that must be met before a merge can proceed.
“Beyond automated checks, lakeFS also supports pull requests, which bring a human into the loop. In agentic workflows, this gives you a way to review and approve what an agent is proposing before it reaches production,” she clarifies.
Who else builds “Git for data” services?
Clearly, other vendors and projects exist in the data versioning market.
Apache Iceberg has functions for branching and tagging data. HPE acquired Pachyderm back in 2023 for its data versioning and pipelines technologies, which serve MLOps teams.
Originally developed by Dremio, Project Nessie is now an open-source data catalog and version control system for data lakes. Data Version Control (DVC) is an open-source data version control infrastructure designed for complex AI operations and big data environments, but now we’ve come full circle as lakeFS acquired the project in late 2025.
In the search for governance automation for agentic AI workloads, lakeFS appears to offer a comprehensive, cohesive set of tools and functions. In the “Git for data” marketplace, a variety of options exist, but lakeFS hasn’t explicitly positioned itself as a carte blanche replacement for similar or related tools.
One thing is certain: The questions of who is feeding what data to which agentic function, when, where, and why are becoming an increasingly pressing issue if we want AI to work correctly.
It’s no secret that generative AI has shifted the operations and business models of companies in nearly every sector. But what if I were to tell you that one day, very soon, we will view these innovations the same way smartphone owners look back on the feature phones of the nineties: early experiments in a journey towards a much more significant technological transformation?
The fact is that AI is advancing faster than any technology in human history; faster than even its creators expected. Today, with many businesses still getting to grips with generative AI, the lens is already moving on to the next world-changing iteration of the technology: agentic AI.
The opportunity of agentic AI for Europe’s enterprises
Already estimated at around $9.14 billion, the agentic AI market is forecast to grow rapidly, at a compound annual growth rate (CAGR) of 40.5%, to reach $139.19 billion by 2034. Europe will be at the forefront of this boom, growing at a 42% CAGR.
The impact of all this investment on enterprises will be profoundly beneficial. According to one study, agentic AI could generate up to $450 billion in economic value through revenue growth and cost savings by 2028.
Today’s stand-alone AI models are rapidly being replaced by complex, multi-agent-driven automation, in which enterprise systems will be able to make decisions and act on them entirely autonomously.
As this happens, today’s stand-alone AI models are rapidly being replaced by complex, multi-agent-driven automation, in which enterprise systems will be able to make decisions and act on them entirely autonomously. I’m not saying that businesses should abandon their current generative AI efforts. They need to start putting the foundations in place for the agentic future now.
Ensuring cloud infrastructure is agentic-ready
Vultr CMO Kevin Cochrane
Agentic AI is a very different proposition to generative AI and demands a different approach within the data center. Rather than serving discrete models, agentic AI infrastructure will need to orchestrate multiple autonomous systems as they interact with one another and with human users.
From a technical perspective, that means balancing both high-performance cloud GPUs and CPUs in an end-to-end AI-optimised stack.
GPUs will be required to run massive LLMs, process data, and generate outputs, while CPUs will need to orchestrate agents and execute the tools and policies that enable them to be autonomous. Modern cloud infrastructure needs to balance these technologies to enable agentic applications to deliver on their considerable promise fully.
However, for European businesses, deploying agentic AI-ready cloud infrastructure comes down to much more than technical capabilities alone. To be fit for purpose, the European cloud offerings of tomorrow will need to be completely different from those currently in play. There must be a decisive break from the past.
Data sovereignty and the need for regional AI infrastructure
Currently around two-thirds of European cloud services are provided by US hyperscalers. This situation is a legacy of a vanished world with a different set of geopolitical, economic, and regulatory conditions to our own. It harks back to a time when storing the data of European citizens and businesses overseas was much less problematic, and when cloud costs were more manageable.
Today, EU businesses need to balance agentic innovation with a renewed focus on regulatory compliance. This is because the EU is mandating measurable standards for data localization, operational control, and legal jurisdiction to address concerns about foreign surveillance risks, extraterritorial jurisdiction, and dependence on a small number of hyperscale providers. Once abstract concepts of data sovereignty have transformed into strict operating principles for European businesses. Cloud infrastructure must be stored locally and in the right operational jurisdiction to avoid foreign data subpoenas.
Addressing exploding cloud costs
A second consideration is cost. Data from Flexera shows that 81% of businesses still cite cost efficiency as their top metric for assessing progress against their cloud goals. Yet according to its research, 76% of large enterprises spend more than $5 million on the cloud each month, and nearly a third complain of wasted cloud spend.
Part of the problem is that enterprises are locked into hyperscaler contracts characterized by opaque pricing and forced service bundling. At just the moment they need to invest in agentic AI infrastructure, enterprises are struggling to fund their core cloud workloads, a situation that’s not helped by the skyrocketing price of CPUs.
As European businesses set out on their agentic AI journeys, the case for moving beyond the hyperscalers could not be more compelling.
Rooting cloud infrastructure in Europe
This is where alternative hyperscalers like Vultr come into their own. From a data sovereignty perspective, we operate nine European cloud data center regions, including Amsterdam, Frankfurt, London, Madrid, Manchester, Paris, Stockholm, Warsaw, and, as of 19 May, Milan (with this launch, we now operate 33 global cloud data center regions).
These are physically isolated data centers that come with geo-fenced data management policies and guarantees that no data will be transferred or processed outside jurisdictional boundaries without explicit consent.
As well as helping European businesses comply with data sovereignty mandates, Vultr helps them avoid the high costs and systemic lock-in associated with traditional hyperscalers. With Vultr’s full-stack AI infrastructure, developers and enterprises can benefit from Vultr’s flagship CPU offering, VX1, which offers 23% better performance and 33% lower cost than comparable hyperscaler compute plans, resulting in up to 82% better price-to-performance.
In addition, Vultr offers European businesses with access to the latest AMD and NVIDIA GPUs for AI and machine learning, high-performance computing, and more, available on demand either as virtual machines, bare metal, or self-service clusters. This is the complete, end-to-end stack of next-generation compute power that businesses will need to thrive in the era of agentic AI.
The dominance of the hyperscalers is less certain than ever as data sovereignty mandates make the case for Europe-based infrastructure.
It’s an exciting time to be in the cloud infrastructure business. The dominance of the hyperscalers is less certain than ever as data sovereignty mandates make the case for Europe-based infrastructure. Meanwhile, businesses are pushing back against the high costs and systemic lock-in that come with hyperscaler offerings. In their place, enterprises can invest in high-performance, cost-effective compute infrastructure that’s open, flexible, and ready for the demands of agentic workloads.
About a year ago, Google demoed a diffusion model at its I/O developer conference, but went quiet about the technology soon after.
On Wednesday, however, Google broke that silence with the launch of DiffusionGemma, an experimental 26B mixture-of-experts model that uses diffusion to generate text 4x faster than its existing Gemma models.
Diffusion has long been the standard for generating images (think Stable Diffusion). Instead of generating one word at a time, models like DiffusionGemma or Inception’s Mercury 2 generate words in parallel.
At first, those blocks of text don’t make sense and seem random. But then, with each new step, the model refines the text and reduces the noise until it becomes the answer you were looking for. If you’ve ever looked at a diffusion image model generate images in real-time, that’s essentially the same process, but for text.
Credit: Google
With each step, the model denoises 256 tokens in parallel, which is why it can be much faster than a traditional autoregressive large language model. It basically iterates on the text with each step until it.
All of these tokens attend to all others, which Google says is especially helpful for use cases such as inline editing, code infilling, working with amino acid sequences, and mathematical graphs.
Credit: Google
Google says DiffusionGemma can produce more than 1,000 tokens per second on a single Nvidia H100. And since the model uses the mixture-of-experts technique, it doesn’t have to keep the full 26 billion parameters in memory; instead, it activates only 3.8 billion during inference. This means it can easily run on a GPU with 18GB of VRAM.
There are some tradeoffs, though. On all benchmarks, the DiffusionGemma model underperforms when compared to Gemma 4 26B A4B. That’s something Google itself acknowledges. There’s no technical reason why a diffusion model couldn’t perform just as well as a more traditional large language model, but the focus here is on speed.
“For applications that demand maximum quality, we recommend deploying standard Gemma 4,” Google says in its announcement.
Credit: Google
Availability
The model is now available on HuggingFace, with Unsloth and other quantizations available for those who want to run it locally using llama.cpp and (soon) similar local inference tools.
Google also worked with Nvidia to optimize the model for its hardware, including high-end GPUs like the GeForce RTX 5090 and 4090, as well as the Nvidia DGX Spark and DGX Station (for those who can afford them). Nvidia NIMs are also available for the model.
“Keep readers reading” is the not-so-simple goal of Medium’s recommendations system. To predict what’s most likely to appeal to a particular reader at any given time, Medium continuously processes user activity signals (stories read, recommendations shown, follows, likes, etc.). It then immediately correlates that with the steady stream of new articles, which is estimated at millions per month.
Smart models and good inference logic are required, but that’s not enough. The data must be stored and retrieved quickly enough to remain relevant while the user is browsing. That’s the job of Medium’s feature store. And getting the data model right started to matter a lot as they scaled to 1M operations per second.
Andréas Saudemont, Medium Principal Software Engineer, recently walked through how the team identified the problem and what they built to fix it. If you’d rather watch than read, you have two options: Watch a short version from Monster Scale Summit or an extended follow-up webinar
The feature store and its role in Medium’s recommendation system
The feature store ties it all together, ingesting user activity and internal events and feeding them to the ML models that power recommendations. It’s what enables customization like the “For You” feed that greets logged-in users.
Each feature is a property of an entity, usually a user or a story. Some are simple and static, like whether a user holds a paid membership. Others capture interaction history: which stories a user has read, what content they’ve recently been shown, etc.
The following diagram shows a highly simplified view of the Medium feature store architecture:
The problem with a relational features data model
When they built their feature store years ago, Medium used relational features for cross-entity relationships. Unlike regular features, a relational feature can have multiple values for a given entity ID. Each value is defined by a relation ID (the ID of the related entity) and a timestamp recording when the event occurred.
For example, a “story users have read” feature is attached to the story entity type. It relates to the user entity type, and its values indicate whether/when a given user has read that story.
Andréas shared the following schema diagram to explain the concept:
Features sit at the center, each attached to an entity type and defined by name, version, and data type. Non-relational features are simply a feature, an entity ID, and a value. Relational features add a relation ID mapping to another entity type, plus the value itself and a timestamp.
This approach proved suboptimal from a data modeling perspective. Since relational features link two entity types, the data ends up split between two tables: one for the entity IDs and one for the values. That means you can’t get both in a single query. The first query retrieves only entity IDs (not their associated values) and relies on ALLOW FILTERING. A second query then runs for each entity ID to fetch its value. “If we have 1000 entity IDs for which we want to fetch values, then we have to run 1000 queries to fetch these values,” Andréas said.
Overrelying on ALLOW FILTERING made things worse. “This is bad,” Andréas said, referring to monitoring data showing that 90% of rows read via these queries were simply discarded. “This is just data that we don’t need. ALLOW_FILTERING should be an escape hatch, not our design pattern.”
“ALLOW_FILTERING should be an escape hatch, not our design pattern.”
The list feature model
So they reinvented their data model and shifted to a list-based feature model. Instead of splitting data across two tables, everything for a given entity lives in one place and is retrieved in a single query.
Like other features, a list feature is defined by its entity type, name, and optional version. What’s different is the value. While a non-relational feature has a single value, such as true or false, a list feature’s value is a collection of items, each containing a value and a timestamp. Item values can be of any data type; the feature store doesn’t enforce consistency within a list.
For example, consider a user’s reading history. The entity is user, the feature name is reading history, the TTL is 6 months. After that TTL is reached, the data is automatically dropped by the database (since older history isn’t useful for recommendations). The list for a given user is a collection of story IDs and the timestamps at which they were read. The same story can appear multiple times, and multiple items can share the same timestamp.
A range of operations need to be supported. Create List and Delete List operations run at most a few times per day. Remove List Items with Value, which lets a reader scrub a specific story from their history so it stops influencing recommendations, runs at 1k-10k per second. Add List Items is higher still: every story read and every thumbnail shown to a user generates an event. Get List Items is the top, at 100k-1M operations per second.
“The Add List Items, and even more the Get List Items operations, are really the reasons why we need an efficient data store.”
“The Add List Items, and even more the Get List Items operations, are really the reasons why we need an efficient data store,” Andréas said.
Multiple items, one timestamp
Beyond raw efficiency, the new data model also had to support multiple items with the same timestamp. When Medium shows a user four story thumbnails simultaneously, all four presentation events share the same timestamp, but have distinct story IDs. If this isn’t handled correctly, primary key collisions occur.
The team’s solution was a single list_items table that stores everything.
The partition key combines feature_key and entity_id, keeping all items for a given list together. All of user 123’s reading history is stored in one partition, retrieved in one query. The clustering key concatenates each item’s timestamp with an MD5 hash of its value. The hash is what makes same-timestamp items with distinct values possible.
Relying on MD5 hashes for uniqueness raises its own set of questions, but in practice, the team hasn’t seen collisions. “The values that we are storing are sufficiently distinct, especially when you add the timestamp into the equation,” Andréas said. The table’s clustering order is set to descending so ScyllaDB can optimize for the typical read pattern (most recent N items) rather than leaving the application to sort afterward.
TTL to control storage costs
Storage cost is controlled entirely through ScyllaDB’s native TTL, with no cleanup logic required. Every row expires automatically based on its own timestamp plus the feature’s TTL duration. “We don’t have anything to do regarding that,” Andréas said. “Any row for which the TTL is expired will be considered deleted by ScyllaDB.”
Storage plateaus for a steady write rate. When a feature is retired, its data drains away on its own. “That’s super useful for controlling our storage and usage costs.”
Implementing the list operations
Add List Items is a logged batch of INSERTs with atomicity guaranteed: all items land or none do. Each row carries its own TTL calculated from its timestamp, so older items expire sooner. Since items almost always carry a current timestamp, new entries append to the top of the partition, which is exactly where reads will look first.
Get List Items runs as a single-partition SELECT with a minimum timestamp and a row limit. “We run the query on a single partition,” Andréas said. “That’s the maximum efficiency that we can have.” The clustering key handles filtering and ordering directly. Post-processing is not required.
Remove List Items with Value is the one operation that couldn’t be reduced to a single query. Because value isn’t part of the primary key, a direct filter isn’t feasible.
A local secondary index built specifically for this case first finds the matching item keys, then a batch DELETE removes them by primary keys.
“Using an index is really faster than a scan because the query is highly selective,” Andréas explained. “We have very few items in a given list that have the same values compared to the total number of items in a list. And thanks to the current structure, using a local secondary index is faster than a global index.”
Andréas shared another example. Starting with the original table partition, the goal is to delete all items with the value “storyC.” Using the local secondary index, the system first identifies the two rows containing that value. It then issues two DELETE statements using the item keys from those rows, which removes them from the list. The final operation, removing all list items, is even more straightforward.
“We can just drop the partition,” Andréas said, “and ScyllaDB does its magic. It just deletes all the rows for that partition, which means that it deletes all the items for the given list. And bonus point: it’s atomic. It’s either completing successfully or not changing anything at all.
ScyllaDB vs. DynamoDB performance
Medium implemented the list operations on top of both ScyllaDB and DynamoDB. The main goal was to benchmark how both databases compared on their actual production data. “Conceptually they are very close,” Andréas noted, “but they have significant differences in how they operate.”
For AddListItems, P50 latencies were low with both databases: ScyllaDB came in under 1.5ms, DynamoDB under 5ms. “DynamoDB is extremely fast, not as fast as ScyllaDB, but extremely fast at sub 5ms latency,” Andréas commented. Things got more interesting at the P95 and P99 latencies. ScyllaDB held steady at around 5-6 ms P95s, while DynamoDB ranged from 13-45 ms. ScyllaDB’s P99s were steady single-digit milliseconds, while DynamoDB’s ranged from 40- 120 ms.
AddListItem latencies: The blue line is DynamoDB; the purple line is ScyllaDB
It was a similar story for GetListItems. At P50, ScyllaDB clocked in at 1 ms, DynamoDB at around 3.5 ms. At P95, ScyllaDB held around 5-6 ms while DynamoDB spiked from 30 – 60ms. And at P99, ScyllaDB remained at ~30ms while DynamoDB ranged from 70 ms all the way up to 220 ms.
GetListItem latencies: The top blue line is DynamoDB; the lower purple line is ScyllaDB
“ScyllaDB is very fast, with very predictable performance, and that’s super important for us.”
One caveat: DynamoDB was running without an extra caching layer. “We expect that could have a significant impact for DynamoDB because of the high cache hit rate that we are seeing on the list,” Andréas said. “But we don’t have the data yet, so we cannot compare them.” His verdict for now: “ScyllaDB is very fast, with very predictable performance, and that’s super important for us.”
Key takeaways
One pleasant side effect of getting the data model right: Medium is now eager to use ScyllaDB for additional feature store workloads. Before, they were holding back because they didn’t want to build on the shaky relational feature foundation.
Reflecting on the path to this point, Andréas left the audience with this parting advice:
“If you have a suboptimal data model, you will have queries that are slow, that will scale badly. And most likely, you won’t be able to optimize that data model. You will have to define a new data model that will be better. So take time to think about your data model before you start the implementation, because once you have production data using your suboptimal data model, it’s too late.”