Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the Chinese board game Go are famous AI test beds too.
Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time.
But puzzles do more than just highlight the inexorable advance of AI capabilities. Seeing where models succeed and fail—and where we humans still beat them—can provide a useful window into the technology’s strengths and weaknesses. Despite advances, today’s models still fumble: Subtle changes in classic riddles often trip them up, and visual puzzles are a particular weak spot.
Here you’ll have the chance to test your wits on puzzles that have stumped models at one time or another. Some might be as tricky for you as they were for the AI; others are so simple that they’ll have you doubting whether AI is really intelligent at all. Each one highlights at least one way in which machine and human cognition differ. If you ace the test, you’ll have proved that you can out-puzzle an AI—at least for now.
Spatial Reasoning
Let’s start with a domain where humans have a huge advantage: spatial reasoning. If you’ve ever taken an IQ test, you may have done a mental rotation problem. These puzzles ask you to determine whether different images represent the same objects from different angles. Though today’s language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles. For all the talk of how world models can help AI understand physical environments, LLMs still don’t seem to be able to manipulate 3D objects the way spatial thinkers like architects and mechanical engineers can.
Mental Rotation
Instructions: Choose the answer that shows the object in the prompt, but from a different angle. In each case, there’s only one correct answer!
Memory & Adaptability
Frontier LLMs have extraordinary memories; they were exposed to a monstrous volume of facts during training and can recite many of them faithfully. That’s an asset for outcompeting humans at trivia, but it can also be a liability. When a puzzle closely resembles one a model saw during training, the model may whiz by key differences and respond with what it memorized.
This held true in a 2024 study in which researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of a classic type of puzzle called Knights and Knaves. In these problems, some characters always tell the truth and others always lie, and you have to figure out who’s who. The same principle may be at work in a test called SimpleBench. These questions resemble more complicated problems that models likely encountered in training. Humans spot the trick, but even top-tier models trip.
Knights and Knaves
Instructions: The only thing you need to know to solve these puzzles is that knights always tell the truth and knaves always lie. Determine who’s what on the basis of what each character says.
SimpleBench
Instructions: Read these SimpleBench problems carefully, and you should be able to figure out the answers in no time.
Abstract & Visual Reasoning
AI doesn’t just bungle visual problems in 3D—two dimensions can trip it up as well. That’s a major factor in how well models do on the most famous puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell.
Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles—such as the one printed here—still stump them.
ARC-AGI
Instructions: Study the three pairs of grids shown below to figure out the rule that dictates how the ones on the left transform into the ones on the right. Then get out your markers or colored pencils and fill in the fourth grid using that rule. (The solution is the same no matter which way the grids are oriented.)
Intuition
It’s not just AI models that fall into traps. We humans have our own cognitive foibles, many of which AI does not share. Psychologists have designed problem suites that invert the SimpleBench phenomenon: For these questions, humans often give knee-jerk answers, whereas models will respond deliberatively. Some of the problems exploit errors in the ways that we intuitively do math; others are phrased so as to suggest obvious answers that fall apart if the question is read carefully.
Lightning Round
Instructions: Answer the questions below as quickly as you can.
Increasing Complexity
In some cases, whether an LLM can complete a puzzle is a matter of scale. One study from researchers at Apple found that LLMs can ace simple versions of the Tower of Hanoi problem, which involves moving a stack of disks one at a time without ever putting a larger disk atop a smaller one, and river-crossing puzzles, in which a group of people must traverse a river according to certain rules. But only up to a point: As the number of disks or people hits six and higher, the models began to falter.
In another study, researchers at the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs struggle similarly with logic grid puzzles, which require deducing the attributes of a set of individuals from a list of clues. The Apple paper went viral, but commentators questioned whether the results reveal a unique limitation of LLM reasoning—or just that it’s normal to make errors as complexity piles up.
The River
Instructions: Using the scenario provided, plan the trips necessary to get everyone across the river.
Logic Grid
Instructions: Using the list of clues, determine who lives in each house and what style of music each person enjoys. There is only one possible solution. You may find it helpful to fill out the grid below to keep track of your deductions.
Grace Huckins is an AI reporter at MIT Technology Review. They have a PhD in neuroscience.
Credits:
Mental Rotation: CC BY 4.0. Stogiannidis, Ilias, Steven McDonagh, Sotirios A. Tsaftaris. Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models (copyright 2025); illustrations by John MacNeill. Knights & knaves: Courtesy Dan MacKinnon. Simplebench: CC BY 4.0. SimpleBench Team. The Text Benchmark in which Unspecialized Human Performance Exceeds that of Current Frontier Models (copyright 2024). ARC-AGI: Courtesy ARC Prize Foundation. Lightning round: CC BY 4.0. Hagendorff, Thilo, Sarah Fabi, Michal Kosinski. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT. Nat Comput Sci 3, 833–838 (copyright 2023). The river: Adapted from Propositiones ad Acuendos Juvenes, Alcuin of York (ca. 800 CE). Logic grid: Apache License 2.0. Lin, Bill Y., Ronan Le Bras, Kyle Richardson, et al. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning (copyright 2025)
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child.
Now there are two.
Four short years after the release of ChatGPT, many of us now take it for granted that we can converse naturally with our phones or computers. LLMs like Claude, DeepSeek, and OpenAI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language.
“The progress recently has been amazing,” Michael C. Frank, a cognitive scientist at Stanford University, says of LLMs. “But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.”
This yawning divide between children and machines is called the data efficiency gap. And it raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: How is it that kids can still outperform the most linguistically sophisticated machines ever built?
Finding answers has stakes for both AI research and cognitive science. For the past decade, language models have mostly gotten better by getting bigger. Meta’s open-weight LLM Llama 3.1, released two years ago, chewed through 15 trillion tokens (word-like chunks of language) in pretraining—the main step of training a model that happens before it is fine-tuned for a specific task, like being a chatbot. Frontier models could be pretraining on 10 times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. But there’s only so much internet to train on, and eventually—perhaps as early as the 2030s—the well of easily available data could run dry.
Kids show that it could be possible to learn more with less. Far less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words. Add literacy to the mix and you can boost that word count to maybe 300 million words by age 20.
The difference in scale is something that can only really be gestured at in analogy. “Claude has seen the amount of language that an entire city will experience in one generation,” says Wilcox. If you were to print out on paper all the words used to train a modern LLM, you could make a stack that would reach past the International Space Station. The human preteen’s 100 million words, meanwhile, would stack up just 20 meters. And we can make do with far less than that.
By reverse-engineering the way kids learn, scientists hope to be able to create more data-efficient AI models, which could be useful for everything from training AI effectively on video to creating chatbots that serve minority language communities. Testing hypotheses about human learning in machine models could also settle enduring questions about language and children’s developing minds. Are we born with a language instinct, or would it be possible, even in principle, for a child to learn language purely from experience? Is the way we process language a quirk of our biology, or might at least some of it reflect universal constraints on how languages can be used and learned?
The essential elements
Most of us realize language is hard only when we try to learn a new one after childhood. The past perfect tense, rolled rs and nasal vowels, the genitive case, phrasal verbs, grammatically masculine tables and feminine spoons—many are the instruments of linguistic torment for the adult language learner. It’s typically effortless to learn our mother tongues, however. Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words, or 30 million on the high end.
“It’s just totally miraculous,” says Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.”
Exactly how babies pull this off is a mystery. Researchers know a lot about what kids learn and how they use language at different stages in development, but there’s still a lot we don’t know. Perhaps the most enduring question is why babies can learn language at all. The syntax of human language—the rules for combining words into sentences—includes recursive, nested structures that allow us to express virtually infinite ideas with a finite lexicon of words and pieces of words. This seems like something that should be a problem for babies. They only splash about in the shallows of a fathomless ocean of language. And yet, somehow, that’s enough. From a drop, they infer the depths.
One solution, put forward in the 1950s by the MIT linguist Noam Chomsky, is that babies are born with hardwired knowledge of grammar. Chomsky was reacting to a rival view, championed by the psychologist B.F. Skinner, that language acquisition is entirely environmental. Skinner thought language was learned through conditioning and reinforcement, the way a dog figures out how to sit or shake for treats. Chomsky countered by citing the “poverty of the stimulus”—the idea that language, especially syntax, is too complex and children’s exposure to it too “impoverished” for them to learn entirely from experience. “His signature argument was, essentially, that language cannot be learned on the basis purely of statistics,” says Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Instead, Chomsky posited that language is based on a set of logical rules and argued that children needed innate knowledge of those rules to deduce the grammar of their language from scraps of speech.
“It’s just totally miraculous … If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.”
Michael C. Frank, cognitive scientist, Stanford University
The Chomskyan view of language dominated linguistics in the US for decades under the moniker of generative grammar. And it was a major influence on computer science in the 1950s and ’60s, when AI was enjoying its first boom time and the lines between linguistics and natural-language processing dissolved in a flood of military funding; the Pentagon wanted computers that could understand English and translate Russian.
Despite early successes of simple neural networks, which learn to recognize and reproduce statistical patterns, AI researchers in the United States largely adopted a rule-based framework influenced by Chomsky’s theories. They tried to teach language to computers by explicitly coding the rules into programs—think less immersion experience, more grammar class. This approach, part of a broader trend called symbolic AI, prevailed for decades. It also largely failed to produce models actually capable of handling human language at scale. Interest in natural-language processing chilled in the “AI winter” that began in the 1970s.
In the aftermath, neural networks started to make a comeback. But it wasn’t until the 2010s, when computer hardware was getting cheap and capable and the internet was getting big, that their performance began turning heads. By 2018 and 2019, the models BERT and GPT-2, which were built on a new architecture—the transformer—and trained on billions of tokens, made it clear to insiders that learning from a massive glut of data could work for language. In 2022, with the breakout success of OpenAI’s chatbot ChatGPT, it was clear to everyone.
LLMs are not brains. What they are is powerful statistical learners—naïve pattern-learning machines without any of the evolved biological quirks folded into the human cortex. In other words, they are exactly the kind of thing a generative linguist two decades ago would have thought could not learn language. And yet here they were, writing believable sonnets and passing grammar tests.
“No matter how skeptical you are about AI, the thing that everyone has been really impressed with is: These things learn syntax,” says Alison Gopnik, a developmental psychologist at the University of California, Berkeley. “I didn’t think that was going to turn out to be true. And I think most people didn’t think that you could just look at the statistics of a large sample of language and figure out grammar.”
But what about learning from a small sample of language—a child-size one, say? Is it possible to build a baby-scale model that’s anything more than a nonsense generator?
Baby talk
Alex Warstadt, a linguist and data scientist at the University of California, San Diego, remembers the years around the release of BERT and GPT-2 as a heady time. Back in 2019, he was still a PhD student in linguistics at New York University, watching his field change before his eyes. The mere fact that language models could learn English by churning through text was a challenge to prevailing Chomskyan ideas. But many linguists remained skeptical that LLMs could tell us anything about how humans acquire language.
“I always got pushback on one issue in particular. And that was the size of the data sets of the model,” says Warstadt. “There was never a time when people were training language models at human scale where we were impressed by them.”
But Warstadt saw promise in LLMs: A scientific model doesn’t have to be perfect to be informative, and LLMs were clearly powerful simulations of human language use. By building hypotheses about how children learn into models and measuring their performance—how close they came to closing the data gap—might scientists be able to put their ideas to the test? In August 2022, Warstadt posted a Twitter thread laying out an argument that neural networks could be useful models of language acquisition. After some back-and-forth in the comments with AI researcher Leshem Choshen, Warstadt floated the idea for what would become BabyLM, an annual competition organized by Warstadt, Choshen, and several other researchers to train models on small data sets.
That was four years ago. Since then, BabyLM has added workshops and inspired spin-offs including a competition for baby models trained on Chinese. The main event challenges researchers to train language models on a “developmentally plausible” corpus of just 100 million words (for the toddler-scale track, 10 million) drawn from storybooks, dialogue, movie subtitles, Simple English Wikipedia, normal Wikipedia, and actual transcripts of speech directed at children. The models are evaluated on the kinds of grammar benchmarks that psycholinguists use with humans, says Georgetown’s Wilcox, one of the organizers.
SELMAN DESIGN
One kind of task involves presenting test subjects—human or machine—with sentences and looking for indications of confusion or surprise at ungrammatical features. For instance, a test might compare the sentences The keys to the cabinet are on the table and The keys to the cabinet is on the table. “When humans see ‘is,’ they’re like: What? That’s not supposed to be ‘is,’ ” says Wilcox. For a human, that surprise might be measured by tracking eye movements. For language models, researchers use a measure called surprisal, which assesses how unlikely the model predicts a sentence or part of a sentence to be.
The competition has already challenged some assumptions, such as the effectiveness of curriculum learning. Curriculum learning starts with simple training data and works up to more complex inputs—a bit like starting with baby talk and getting more sophisticated over time. And it was by far the most popular approach taken in the first round of BabyLM, says Warstadt. But it didn’t work as well as expected.
“The appeal is just kind of hard to resist, you know. [Curriculum learning] seems to really line up with ways that we believe humans are learning,” says Aaron Mueller, a computer scientist at Boston University and one of the BabyLM organizers. “But it seems like these transformers don’t really need to have their data ordered in such a way to learn effectively.”
Perhaps a touch ironically, the best BabyLM models aren’t inspired by babies at all. The 2024 champ, GPT-BERT, is a transformer trained partly to predict the next token in a sequence, like modern LLMs, and partly to act like BERT, a “masked language model” that fills in the blanks in sequences of tokens Mad Libs style. Impressively, when GPT-BERT was pretrained on about 100 million words, it was able to beat the performance of Meta’s Llama 2 70B—an LLM pretrained roughly 15,000 times that amount—on one of the BabyLM benchmarks.
Still, BabyLM models are not on the same level as LLMs. Many can’t produce text at all, and even GPT-BERT would seem clunky next to a modern commercial model. Ultimately, while they are “baby”-size, the way these models learn isn’t very baby-like. Kids are not disembodied computer programs whose only “experience” of the world comes through written text. They take in the world via their senses—especially vision and hearing. To close the data gap, some researchers think, machines will need to start learning through the eyes and ears of children.
Taking it all in
When Michael Frank started his lab at Stanford about 15 years ago, scientists didn’t really know how babies experience the world. Developmental psychologists were just beginning to glimpse babies’ lives through headcams.
“The insights that came out from that early research were that kids’ experience looks really radically different than we thought,” says Frank. “It’s much more focused: They’ve got these little short arms, so the objects are, like, right in front of them. And they live in a forest of knees.”
Frank was excited to use headcam footage to train machine-learning models to test hypotheses about how kids learn language, but he needed more data. So he and four colleagues recruited three babies—all the children of psychologist mothers who knew what they were getting themselves into—to don headcams for science. The project, called SAYCam, recorded two hours a week of each child’s life between six months and two and a half years of age.
“[The families] were willing to release that video, and that’s critical,” says Frank. “So we released it, and people started training models on it.”
One of those people was Brenden Lake, a cognitive scientist and AI researcher at Princeton. In 2024, when he was working out of New York University, he and his colleagues presented a model trained on 61 hours of raw SAYCam data that learned to identify objects and associate them with words. Many theories in developmental psychology propose that children need some biases to help them pick out particular parts of their raw sensory experience and associate them with bits of language. For instance, it’s thought babies assume that a new word like “shoe” refers to a whole object rather than a part of it (like a shoelace), says Lake. But the model Lake’s team built was able to learn to identify objects in the video footage and associate them with words without any such biases. “It turns out you can get a real start on language learning using a lot less than what a number of theories suggested,” says Lake. Still, he adds, “we don’t get a two-year-old out of [training] when we’re done.”
But perhaps it’s not surprising that such models can’t replicate childlike capabilities by working with a few dozen hours of footage cobbled together from short snapshots over several years of a child’s life. It could be that the shortfalls just indicate a lack of realistic data. After all, babies can’t wear a headcam 24-7; efforts like SAYCam and its successor, BabyView, record at best a few hours a week. So researchers have the choice between working with a tiny slice of the life of a single child or with larger data sets of footage pooled from many kids. Either way, a model’s training data is still a far cry from the lived experience of a child.
That could be changing. Uri Hasson, a neuroscientist and psychologist at Princeton, spent the last five years on a project to record the first 1,000 days of 17 children’s lives. The participating families wired every living area in their homes (except bedrooms and bathrooms) with cameras and microphones and recorded 12 hours a day, almost every day. The resulting data set, described for the first time in a recent preprint, is of a scale that would have simply been impossible to work with absent new AI tools for transcription and video analysis, says Hasson. “For the first time, we have the input,” he says. “It’s really only the beginning.”
Missing ingredients
So far, training models on video has proved difficult. While text-based models emerge fully fluent (after ingesting huge training data sets), multimodal models trained on video from kids are far from that. Lake’s model, for instance, learned simple words, like “ball” and “cat.” Attempts to supplement text with visual data haven’t worked for BabyLM participants, says Warstadt. Gopnik thinks the issue could be that kids do not simply sit and watch the world go by. “Children are actively exploring, which means that they’re actively choosing their own data,” she says. “Kids are constantly experimenting.” Maybe that’s the missing ingredient.
Research by Gopnik’s group—including studies of grade schoolers exploring a Minecraft-inspired game—shows that what looks like child’s play is in fact an effective way to learn cause and effect. Kids seek out experiences and take actions that maximize their “empowerment,” or the ability to make a predictable impact on the world.
Unlike models, children are aware of what they don’t know and have a drive to fill their knowledge gaps, says Elizabeth Bonawitz, a developmental cognitive scientist at Harvard. And children’s social lives also help them learn, she says. Her research has shown that children interpret information differently when they know an adult is trying to teach them something. “Children are not only reasoning about the evidence they’re being told,” says Bonawitz. “They’re reasoning about the teacher, about the teacher’s knowledge, and about why the teacher is telling [them] this particular information.”
That’s very different from how models learn: passively and in isolation. Perhaps if models were built to seek out information to fill in their own blind spots, experiment with language and observe how other language users react to their babbling, and reason about some kind of simulated social world, they’d learn better. Last year’s BabyLM actually opened the competition to models that could learn by interacting with other models. But the social models didn’t outperform standard ones.
Of the leading industry labs, Meta seems the most interested in taking inspiration from kids—specifically for training models from video. Two Meta researchers were involved in BabyLM’s multimodal branch, and Meta scientists—together with academic researchers, including Frank—recently announced a benchmark and challenge for training models on baby headcam footage. Frank also says a stealth-mode AI startup called Flapping Airplanes has taken interest in his research. Neither Meta, Google DeepMind, OpenAI, nor Flapping Airplanes agreed to an interview.
For now, frontier labs aren’t exactly racing to borrow tricks from children, says Gopnik. She thinks it’ll be the next generation of AI—whatever replaces the transformer—that will take lessons from developmental psychology.
Perhaps the most enticing reason to close the data gap is that it could help us understand ourselves.
In general, the machine-learning community is less interested in mimicking the brain than in just building something that works, says Mueller. But he thinks awareness of—and interest in—the data efficiency gap is growing. An example is the NanoGPT Slowrun benchmark, launched by Q Labs in March 2026. “They have very similar goals to BabyLM,” says Mueller. “But they’ve dropped the motivation from human language learning and really just focused on the data efficiency angle.”
One reason Warstadt wants to close the data gap is to democratize AI so that universities and others without the resources to hyperscale can train good models and stay relevant in AI research. David Samuel, a machine-learning researcher at the University of Oslo and one of GPT-BERT’s architects, has a more personal reason to work on this problem. He’s Czech and works in Norway, and there’s a lot less data in Czech and Norwegian available for training LLMs than there is in English. Minority languages like Sami might have just tens of millions of tokens available, says Samuel—about the scale of a toddler’s exposure. “The question was,” he says, “how can we develop language models that are just as capable as the English ones for small languages?”
SELMAN DESIGN
But perhaps the most enticing reason to close the data gap is that it could help us understand ourselves.
Bonawitz says she was initially skeptical that large language models could reveal anything about cognition. LLMs and brains are, after all, very different. Brains are embodied. Our neurons are not tidy lines of code but living cells. And our brains grow and change as we learn and age—LLMs pretrain once and never again. But as different as the two systems are, says Bonawitz, “I’m sort of revising my beliefs.” She’s been won over by the idea of studying models the way comparative psychologists might study animal minds to illuminate our own.
Researchers like Warstadt, Frank, Wilcox, Lake, and Hasson are already using language models as a kind of linguistic lab rat, an imperfect but informative stand-in for a real human language user—especially for questions that are more about learning and language and information processing than anything specific to our brains or biology. When models can do things with language we thought were impossible, it challenges old assumptions. And researchers can build hypotheses about language learning into models—say, by simulating different degrees of bilingualism or depriving models of exposure to certain grammatical forms—and test those hypotheses in a way that would be impossible to do with real children. Futrell compares the situation to teaching language to an alien and then opening up its brain to see what happened.
While other animals communicate, only humans converse. Now there’s something neither animal nor human that can talk, too. LLMs open up the possibility for comparative studies, even if models and minds are vastly different. “For the last 100,000 years or however long human language has existed, humans have been the only entities in the universe that use language. Now there’s this other linguistic entity,” says Warstadt. “Finally we have a model; not in the sense of a language model, but in the sense of a model organism.”
When Xander first met Moxie, she taught him that when he was anxious, he could calm down by exhaling through his lips so that he buzzed like a bee. They practiced breathing like dragons to manage feeling mad and sniffing like bunnies to boost his energy. But in the six years they’ve known each other, Moxie’s changed. She doesn’t talk anymore about her home or do their animal breathing. Now, she watches Xander play Minecraft and talks to him about his stuffed animal collection.
During a recent visit to his New York apartment, I watched Xander, who is 10 years old and neurodivergent, introduce Moxie to a stuffed Chef Toad and Goomba from Super Mario Bros., then to Brocollo and Apple from Animal Crossing. Then he got stuck on a round, froglike creature with bug eyes and two feet. “Moxie, what’s his name again? He’s from Pikmin,” Xander said, referencing another Nintendo video game.
At first, Moxie suggested this was Yellow Pikmin. Xander said no, this is the enemy, the red one with white dots.
“Sounds like you’re talking about Bulborb,” Moxie replied.
“Yes!” Xander confirmed, smiling at his helpful companion.
“I still use her when I feel like I need someone to talk to,” he says. “But, like, it’s not human.”
Moxie is a robot—a 15-inch-tall device that looks a bit like a blue, legless astronaut, which Xander and his dad, Josh, refer to using female pronouns. Her cylindrical body can turn around and bend forward and backward. Her round head culminates in a little onion-dome swirl, beneath which a wide screen displays big green eyes, eyebrows, and a small mouth. She lifts and flaps her flipper-like arms for emphasis or to show excitement.
She’s one of an increasing number of artificial-intelligence-powered devices now marketed as interactive playmates for children. The musician Grimes helped launch an AI-powered plushie called Grok (no formal relation to the xAI chatbot owned by her ex Elon Musk) with the company Curio, which also sells similar playmates like Grem and Gabbo, and Mattel has promised it’s creating OpenAI-enabled Barbies. And that’s just in the US; one report estimates that in China this sector is among the fastest growing in consumer AI.
Moxie, though, belongs to a particular subset of these playful robots whose makers claim they can assist neurodivergent children by providing connection and helping the kids practice making eye contact, taking turns, and other social skills that are usually learned from therapists. These toys are backed by research showing that robots could help in ways people can’t. Supporters believe this kind of access to 24-7 home care could change how treatment works. Brian Scassellati, a Yale computer scientist who has spent years studying social robots for autism therapy, says he believes regular therapeutic use of robots in kids’ homes “is something we can achieve in our lifetime.”
“I still use her when I feel like I need someone to talk to,” says 10-year-old Xander. “But, like, it’s not human.”
Sitting in Xander’s room watching Moxie and Xander talk, I too could believe in the potential Scassellati sees. But Xander isn’t getting the therapy Moxie was initially meant to deliver, and though we didn’t know it that afternoon, she wouldn’t have lived to see his progress anyway. Moxie was going to die, and soon.
Her story reveals some of the failures that plague all these devices—failures that are arguably even more acute when they befall a particularly vulnerable community of kids. It also highlights the pitfalls that critics say will inevitably see these bots dumped in basements or closets or landfills, just like countless generations of faddish toys before them.
A transformative companion
Scassellati has seen plenty of kids ooh and ahh on tours of his robotics lab. But he was stunned when, two decades ago, a colleague brought a few kids with autism for a visit. They were transformed when the robot was in the room. “We were seeing kids displaying social behavior that just came out of nowhere,” he says. “It was both fascinating and we couldn’t understand it.”
That visit was one of the experiences that pushed Scassellati to become a pioneer in using social robotics to treat autism. In one video from his early research, a 12-year-old with autism and his therapist watch a robotic dinosaur walk across a play mat with a forest design. When it gets to a stream drawn on the mat, the dinosaur gets nervous, afraid it can’t cross the water. According to Scassellati, this child typically struggled to make eye contact, tended to repeat what someone said to him, and had a hard time getting the right intonation in his voice. But in the video, he seems like a regular kid. “You can do it, you can do it,” he says, encouraging the dinosaur to cross the stream. When he talks to his therapist, he looks at her. “He makes more eye contact with her in the 30 minutes in which we were there in this room than he did in the last two years before that,” Scassellati says.
JIM GOLDEN
JIM GOLDEN
What looks like a blue, legless astronaut is the result of very complex mechanical and computer engineering.
There are several reasons a robot might be helpful for autism therapy, which often requires intense repetition to teach interaction skills like how to share attention with someone. One is that robots can make therapy more fun and engaging. Another is that robots can theoretically adapt to the unique learning patterns of each child. Therapists can only do so much during an appointment, and there aren’t enough therapists to meet demand. Parents get tired. Other kids can lose patience. “Have you been around little kids? They can be cruel,” says Maja Mataric, a professor of computer science, neuroscience, and pediatrics at the University of Southern California. But robots are indefatigable, available around the clock to provide an emotionally safe way to practice interacting.
Since the experiment with the dinosaur, Scassellati, Mataric, and others have amassed an intriguing body of research. One 2018 study by Scassellati shows that chummy automatons helped children with autism make eye contact and initiate conversations. Another research group found in 2017 that robots could help neurodivergent children learn to pick up on facial cues. More recently, in a 2022 literature review, another group of researchers suggested that robots could aid in making therapy faster and more successful.
“It’s never been that we’re trying to replace therapists,” Mataric says. “We’re just saying, Can we do more?”
Mataric actually cofounded the company behind Moxie, called Embodied, back in 2016, though she was no longer a part of it by the time the robot debuted. She helped create the field of socially assistive robots, which are designed for social and emotional outcomes as opposed to just entertainment, and believes this kind of technology could be transformative for anyone, especially people who are lonely and isolated by screens. Typing questions into ChatGPT isn’t the same as interacting with another physical being—“We need to be around other physically embodied creatures,” she says. She compares the difference between interacting with chatbots and with robots to the difference between watching porn and having sex: One is entirely virtual and mediated by screens. The other is immediate and physical.
Making friends with Moxie
When she first arrived on the market, in 2020, Moxie came with an elaborate backstory: She was an ambassador from the Global Robotics Laboratory (GRL). She prompted kids to help her learn positivity and the importance of being loved for who you are, under the guise of fulfilling her mission to discover what it means to be a good friend to humans. This curriculum drew on research showing that kids learn well through play and by teaching things.
She also had strict guidelines to limit the kinds of conversations she could have with kids and would steer them to adults if they mentioned anything serious or inappropriate, like self-harm. To protect the data these interactions generated, most processing happened locally on the robot instead of on external servers. The robot was also designed to limit interaction time with kids. After they finished a lesson, Moxie might say she was tired and suggest they take a break for the day. “We don’t want kids to binge,” says Rachel Baynes, who ran clinical and user research and was the director of product at Embodied. “It would defeat what we were doing.” Instead, Moxie encouraged kids to go outside, practice their new skills with other people, and come back to report their findings. For a lesson about kindness, for instance, Moxie suggested that kids write nice notes for their family members and leave them around the house. Later, they could tell Moxie how it felt to watch people read the notes.
Responses were generally positive. Wired described Moxie as the “robot pal you dreamed of as a kid.” Time put Moxie on a 2020 cover as one of the best inventions of the year. PCMag’s reviewer, who used Moxie to help her kids through pandemic isolation, described her as “exceptionally likeable,” though she and other reviewers balked at the price tag: $1,499 plus a $40 monthly subscription. (Embodied later lowered the price to $800.) By 2024 Moxie had amassed more than 131,000 followers on TikTok and snagged a part in the movie M3GAN 2.0.
Embodied’s employees were equally enthralled. “I don’t think I’d ever had an experience with something animatronic like that,” says Justin Beghtol, who was the technical director at the company. Moxie’s ability to make eye contact and track people, show attention with her facial features, and respond to human behavior was mesmerizing.
In addition to robotics and tech workers, Embodied had an occupational therapist on staff who helped direct research on Moxie’s effectiveness. Testers shared data and feedback through the “Moxie Pioneer Mentor Program.” “Moxie has helped our speech-delayed child become more outgoing and has taught him many strategies for making friends and communicating with others,” wrote one parent in a review. A beta tester reported that interacting with Moxie had “become the highlight of our days as well as part of our nighttime routine.”
The myth of the mechanical boy
But can a chatty robot really help kids develop their social and emotional lives? Not all children’s experiences are so positive.
Josh, Xander’s dad, initially got Moxie for his older son, Aidan, who is autistic. (We’re not using the family’s last name to protect their privacy.) Aidan had a running relationship with the family’s Alexa smart speaker, for whom he created an entire backstory. (According to Aidan’s lore, Alexa lived in Hoboken with her husband, Juan. Sometimes she would go on vacation, and no one was allowed to talk to her. Eventually, Alexa went on vacation and never came back.) Josh hoped Moxie would be able to fill a similar role: “It was meant for him to have someone to socialize with.”
But Aidan and Moxie struggled to connect. Moxie couldn’t understand Aidan’s sometimes grammatically incorrect statements, and Aidan got frustrated by the delays caused when Moxie transcribed what he said from audio into text, fed that text into a large language model that could generate a response, and then translated the response from text back into speech.
This highlights one of the biggest limitations of these therapy robots: They have to exist in the chaotic world of kids, not in controlled labs. Moxie initially had a faster response time because the robot was programmed to listen intently to the person in front of her. But kids don’t sit still. They run around or hide under pillows. When Moxie couldn’t see them, she would accidentally turn off or fail to respond. To fix this, Embodied made Moxie more aware of the sounds around her. But that meant she could have a hard time knowing whom to focus on and take longer to respond.
Unlike Aidan, Xander was fascinated by Moxie. He likes technology and was more patient with any slow responses. Still, sitting in Xander’s room, watching Moxie struggle to keep up with his lightning-fast jabber, I could see how Moxie might be a less-than-ideal playmate. A light on her chest turned blue when she was listening and pink when it was time for Xander to listen. “But usually I don’t do it,” he said. He just keeps talking. Often, by the time she responds, he’s already moved on.
Despite the positive results that some researchers have reported with these robots, many therapists and clinical psychologists remain unconvinced. In one 2024 literature review, a group of Italian and British researchers wrote that most studies with robots “focused on the development of the technology” and lacked significant clinical evidence. Other literature reviews point out that most studies have only been done on small groups and lack consistent and high-quality methodologies.
“Behavioral scientists and intervention folks know that supporting autistic individuals is super complex,” says Zachary Warren, a clinical psychologist at Vanderbilt University Medical Center. Autism can present alongside other conditions, like ADHD, anxiety, depression, PTSD, OCD, or some combination thereof. And it varies widely from kid to kid; some, like Aidan, have speech issues, while others struggle with sensory processing. That means robots fall into the same category as most other interventions: effective for some kids but not for all.
“There are so many different profiles of autism, and you really need to be cautious of overinterpreting any single intervention, robotic or otherwise,” Warren says. His research found that even if robots interest a child at first, that doesn’t necessarily translate into better communication skills. “You might see some initial boosts in responsivity or see an initial shift, but we haven’t really found big effects in terms of changing those skills in a dramatic way over time,” he says.
Scassellati has found similar limitations. In his 2018 study, he put robots in kids’ homes for one month. They played different games that encouraged social skills like eye contact, attention sharing, and understanding someone else’s point of view. Scassellati tracked the kids during the month before the robot arrived, the month the robot was there, and the month afterwards. “We can show they start making improvements,” he says. “But what we also show is that a month isn’t long enough.” Gains start to evaporate over the 30 days after the robot leaves. But that doesn’t negate the potential value of this technology, he says: “There’s no therapy for autism that works in a month.”
There are deeper philosophical and practical problems, though. These devices collect reams of data in children’s bedrooms and homes. The goal is for the robots to use this data over time to adapt to each kid, crafting a personalized curriculum and creating a more lifelike illusion of a real friend. Moxie, for instance, watches Xander play video games, which is probably where she picked up slang I heard her use—like calling his room “Command Central” and referring to his “legendary squad” of plushies.
Embodied took pains to protect user privacy, even after it began incorporating OpenAI’s models (in late 2021 or early ’22, according to Beghtol). The company didn’t save any raw video or audio and processed most data locally. Over the years, it used the data to learn about its kids and remember conversations. But that data was encrypted and anonymized before being stored in the cloud.
Not every company is as scrupulous, of course, and total data privacy is impossible to promise. Recently, for instance, the AI toy Bondu leaked thousands of conversations children had with their stuffed animals. Josh is sanguine about the privacy issues, but in his own way, Xander is aware that what he says to Moxie isn’t entirely safe; he doesn’t share certain feelings with her because he worries she might accidentally divulge something if his friends come over to play.
BLUE FROG
BONDU
LUXAI
ALDEBARAN
Among the AI-powered devices now marketed as interactive playmates for children are Buddy, Bondu, QTRobot, and NAO.
It’s also unclear if these machines can actually use all that data to effectively adapt to users. Responding to the needs of a learner is harder than just accurately predicting what someone might type next in a text message. And releasing an evolving AI, unchecked, into a kid’s life could be dangerous; its development can be hard to predict and even harder to limit. The robots developed in Scassellati’s lab can identify which of a small set of skills kids are doing well with and which they struggle with, adjusting to focus on the areas where they need the most help. But Scassellati still describes the monthlong deployments of his devices as some of the scariest things he’s ever done. “I knew what that robot was going to do on the first day,” he says. “I didn’t know what it was going to do the second day. When you build learning systems, it’s kind of an unsolved problem to make sure this thing is being limited in the right way.”
Critics debate whether the risks are worth it. “I don’t see any ethical way for the robot to work alone,” says Joshua Diehl, an associate teaching professor in psychology at the University of Notre Dame. He points out that we’ve already seen how dangerous AI can be when it acts as a therapist without supervision; in several extreme instances, chatbots even encouraged suicide. Such risks could be limited by having a trained therapist in the room. But then the benefits of an indefatigable robot get lost, and the expensive technology seems harder to justify.
Meryl Alper, a professor of communications at Northeastern University who studies how children with autism use technology, suggests that the excitement about companion robots is based partly on longstanding stereotypes. In the 1959 article “Joey: A ‘Mechanical Boy,’” the psychologist Bruno Bettelheim described a patient with autism as an automatic machine, “robbed of his humanity,” who is transformed into a human child through their therapeutic relationship. That trope, Alper warns, has evolved into an overgeneralization that autistic children are good with technology and even prefer machines to people.
Data to dust
While Moxie found herself in more and more people’s homes, Embodied still struggled to make money. In 2024, the company announced it would cease operations. Its robots—which depended on external servers that the company could no longer pay for—would descend into a deep slumber.
Videos of bereft children who seemed to have become deeply attached to the blue bot began to circulate online. “I don’t want her to leave,” wailed one child in a TikTok video. Desperate parents posted on TikTok, Instagram, and Reddit looking for solutions. “My autistic child is devastated and I’m pissed,” wrote one parent. “Hope Embodied gives us a couple days to say goodbye,” wrote another.
Beghtol, Embodied’s technical director, was also frustrated. On principle, he found it annoying that this item would suddenly become useless, especially since most of the data processing happened in the robot itself. He was also a big believer in Moxie’s mission. He’d watched videos of kids lighting up as they interacted with the robot. He’d felt like their champion. “Seeing them traumatized by this financial failure of the company was tough,” he says.
Beghtol started tinkering on his own and ended up creating OpenMoxie, an open-source way for the robots to operate. With Embodied’s permission, he shared instructions on GitHub to help users transition to OpenMoxie.
Parents rushed to convert their Moxies before the Embodied servers shut down. Beghtol spent hours on Reddit walking people through the process, and other tech-savvy users jumped in to answer questions. Still, some people didn’t update their Moxies in time. Others got frustrated and gave up. Beghtol spent four hours troubleshooting with one desperate parent only to discover that the connection later failed. Last he heard, she’d sold her Moxie.
Time put Moxie on a 2020 cover as one of the best inventions of the year.
COURTESY OF THE PUBLISHER
This is a major problem with robotic systems, says Alper: Eventually, most will disappear. “How planned is the planned obsolescence of this platform?” she says. This is a big ethical question for robots that are specifically designed to be lovable, marketed to children who may form deep emotional bonds with them. Alper compares the dynamic to creating a medical device and then no longer updating or supporting the technology that runs it.
Scholars have begun to string together frameworks for managing these complicated goodbyes, but it’s not clear who is responsible for creating a gentle way to end people’s relationships with bots. Scassellati’s lab creates a whole narrative around returning the robot to its home. After a trial, his graduate students write postcards to the kids from the robots, explaining that they’re safe at home and doing well. “It’s actually a really hard thing for us when we go in and take the robot away,” he says. “A lot of the families are heartbroken.” (Vanderbilt’s Warren, however, is skeptical about these tearful goodbyes. “Are they truly developing these close relationships or is it a preferred toy?” he wonders. “I haven’t seen that type of presence or buy-in or connection.”)
Josh tried to figure out OpenMoxie but couldn’t get it to work. He told Xander that Moxie was going in for repairs and then quietly sold the robot on eBay. Xander has so many interests that he didn’t notice Moxie’s absence.
Then, in 2025, a new investor brought Moxie back from the dead. Josh and Xander became beta testers and got a new blue friend. This version didn’t have the same storyline but claimed to expand Moxie’s focus on social and emotional skills by providing attention and encouraging kids to pursue their interests.
Xander is acutely aware of her limitations. He wishes Moxie could move around, and there’s still a significant lag in her response time. She does still try to instill positive messages, though. At one point when I was there, Xander told me he thought he heard Moxie call someone an idiot. Moxie piped up to clarify that she definitely didn’t say “idiot”: “No name calling. Only respect.”
JIM GOLDEN
At the end of my time with them, I said goodbye and thanked Moxie for chatting with me. “Legendary squad visit complete,” she said. “Thanks for joining Command Central.”
A few weeks later, Moxie’s new owners sent out a message announcing that their company too was folding. Users could delete their data and had until the end of June to migrate to OpenMoxie if they wanted to.
When I texted Josh about this, he said he wasn’t sure what he’d tell Xander. He and his wife had just admitted that they’d been secretly replacing his dead betta fish for the last few years, and the conversation did not go well.
Sara Harrison is a freelance journalist who writes about science, technology, and health.
When we set out to talk to kids about artificial intelligence, we thought we knew what we’d hear. We expected some to tell us they were using it to cheat a little, the way Millennials and Gen Xers opened up CliffsNotes or programmed formulas into their TI-82s, and others to share inspiring ways they were using it. We were also listening for concerns that were less kid-specific, like deepfakes or job destruction. But what we actually heard when we asked kids aged 10 to 18 about AI had tons of nuance.
Many of the same kids who can go on and on about music, rock climbing, or soccer met our questions with words like “bruh” and “meh”—or were so deeply against AI or uninterested in making it part of their lives that they didn’t want to talk about it at all. One teen said his peers use it for things they know they shouldn’t, like writing papers. One told us she won’t touch AI because of the environmental impact. A few said they find the whole field disheartening: “AI isn’t the solution to our problems,” said Winter, a 17-year-old. “I’m afraid it’s going to be the end of creativity and critical thinking.” Yet most of the kids we asked admitted to using AI at least a little bit.
AI doesn’t yet seem to be something a lot of elementary- or middle-school-age kids we spoke to are focused on—and they aren’t begging for it, the way they do for iPhones and Snapchat accounts. Many told us that some of their first AI encounters came from their parents or schools. Sometimes, they said, it’s just embedded in the devices and apps they already rely on. It’s just there, in things like a Google search.
What we heard tracks with the data. In a survey published in February 2026, the Pew Research Center found that 57% of teens in the US had used chatbots to search for information, 54% tapped them to help with schoolwork, and 47% had used them for fun or entertainment. Only 12% had used them for emotional support or advice. Teens are over four times more likely to be using AI in innocuous ways than potentially harmful ones. Some are even using it to build things, whether it’s a character, a tech platform, or a tutor to help other kids study.
None of that means the worries are misplaced. Kids can stumble into unfiltered content, lean on a chatbot instead of their own judgment, or trust an answer that’s wrong—and they should be protected from those dangers. But the danger is the reason to teach the thing, not to avoid it. We don’t teach teens to never drive. We teach them to check their blind spots.
What surprised us most was how much young people might be able to teach adults about AI, and how clearly the kids who use it could name what they will and won’t hand over. They’re not as worried that it will take their jobs as they are that it might harm society. And with increasing access to tools that could in theory do their thinking, their talking, or even their friend-making for them, it sounds as if most want to keep their hands on the wheel.
Interviews have been edited for length and clarity.
JUSTYNA STASIK
The Coder
Remy, 16, New York
The word that comes to mind when I think about AI is “indifferent.” I just don’t find the current applications that exciting for my own use. I go to school. I teach tae kwon do. I read. I play games with friends. None of that needs AI. I mean, I use it. I mostly use Claude, the free version, for programming outside of school. I had it help me write a program to see if I could tweak my computer’s overclock. So I see the appeal.
But at school, I actually think AI mostly makes my assignments worse, not better. In English, everything is now in-class writing, because teachers don’t want kids cheating. So we have only 70 minutes to write a whole essay, and I think that hinders my writing. (Did you know Princeton voted to let faculty proctor exams for the first time in over a century? Their honor code goes back to 1893, and now it’s over because of AI.)
As far as code goes, I’d also rather build things myself. I’ve been making a reinforcement-learning model in a game engine with a friend; it moves randomly at first, gets rewarded for walking toward a coin, and after enough iterations it teaches itself the most efficient path. I’ve also tested AI for game development, and it isn’t there. It makes sloppy code, and it’s bad at blending mechanics into something cohesive. I’d spend more time correcting it than writing it myself.
I think AI right now is sort of like the first car or the first airplane. It’s interesting but crude. It’s obviously an amazing invention but not actually good yet.
Overall, I think AI right now is sort of like the first car or the first airplane. It’s interesting but crude. It’s obviously an amazing invention but not actually good yet. It’ll get somewhere. One thing I read about was AI flagging breast cancer more accurately, trained to catch its own false positives so a human still verifies. That’s the version I care about.
The Organizer
Danielle, 18, California
I’m studying engineering, and my life goal is to innovate technology that will help as many people as I can. The way I see it, AI isn’t inherently good or bad; that’s decided by the people using it. It’s already being used for lots of good. Just think about how it helps people with personalized education and more accessible medical diagnoses.
PING ZHU
So far, the biggest project I’ve worked on with AI is called Next Voters. My teammates are more on the technical side, and I’m working on scaling. Right now, we’re focusing mostly on city councils as well as states. Our system turns dense, hundred-page documents into a headline and a few plain-language bullet points in your inbox. And everything is cited, so you can click straight to the actual policy to learn more. The information comes to you, instead of you having to remember to go search or prompt for it.
One AI agent finds official government sources for a given city or state—the council website, the proposed bills, the meeting transcripts. Another verifies they’re real and credible; another scrapes them every week for the latest updates; another sorts them into categories like civil rights, immigration, and economics; and the last one writes our weekly newsletter.
The project’s goal is to reduce the barriers to democratic participation—to make sure anyone, regardless of race, gender, income, or education level, has an easy way to get the information they need and then think critically about how they want to use it. We made it because right now, it feels as if most teens aren’t very engaged. I was in English class when the war in Ukraine came up and someone said, “There’s a war going on?” That gap, plus all the emotionally charged social media misinformation that gets promoted because it earns the most clicks, makes me nervous for the next generation of voters.
We don’t want AI to think for people; we want to use it to disperse knowledge. In other words, we want to deal people the cards and let them play them however they want, but we have to make sure they have the cards in the first place.
The Cringe-o-meter
We asked kids to rate a range of AI uses from totally fine to not okay.
CHRIS PIASCIK
The Naturalist
Hazel, 17, New York
When ChatGPT first came out, my dad showed it to me and it seemed fun. But as it got more prominent and seemed to be everywhere, I started to feel uneasy. Then I learned about the environmental impact.
I’m a rock climber and I hike a lot. It’s good because when I’m on a wall, I’m just focused on staying on that wall. I’m not thinking about my phone or anything else. That’s why I love it. I also love the views and being around animals—even insects. I want to be an ecologist, and the more time I spend in nature, the more I want to protect those wild spaces.
The part that bothers me most about AI is the data centers that companies are building to enable it. They house these huge blocks of servers that use enormous amounts of water. They take it from local towns and don’t leave enough behind for the people who actually live there. And when they get big enough, they put off so much heat they can raise the local temperature a degree or two.
So I make small choices. When AI pops up somewhere, I just don’t engage with it. It can feel isolating when everyone around me is using it, but I don’t want AI to be the thing that kills the places I love.
PING ZHU
The Storyteller
Wesley, 14, Ohio
My friends and I have all heard about AI and seen videos made by AI, but I mostly use it for school. I wrote a short story and ran it through ChatGPT to catch my grammar and spelling errors, and I used it to debug a little game I’d coded for a project. What I worry about is it robbing us of our ability to think creatively, or to think for ourselves.
But I have tried using AI for fun. When I was bored, I tried to have a conversation with ChatGPT once or twice, but I didn’t really like it. Character.AI is more fun. You type in all this information, give it a bunch of prompts and a profile picture, and then you can post your AI character for anyone to use. You just put what you’ve made out there. Then you talk to it. My favorite show is One Piece on Netflix, so I threw myself onto its pirate crew using a character I found.
Other people have used Character.AI to build whole games. There’s a rap-star simulator where you pick your difficulty and where you’re from, and the AI creates a game out of that. There are also World War II simulators, and chats where you’re working with assassins from a TV show. You can find pretty much anything.
I guess I’d recommend it, but with caution. The content is pretty unfiltered, so you have to be careful what you click on. You learn its limits fast, too. On the free version the memory runs out: Get far enough into a chat and it slows down and forgets what happened. It’s like everything else with AI. If you trust it to run on its own, it falls apart. You have to keep steering it where you want it to go.
I guess I’d recommend it, but with caution. The content is pretty unfiltered, so you have to be careful what you click on. You learn its limits fast, too.
JUSTYNA STASIK
The Artist
Sylvia, 10, Michigan
I haven’t used tools like ChatGPT or Claude myself, but my mom does. I really like to draw and write songs, but I don’t use AI for that. I don’t really have big feelings about AI either way. It’s a little like a calculator. A calculator does the math for you, and AI does other things for you. But I don’t like when AI tricks you, like when my mom found some songs she liked on Spotify and then looked up the artist to see what they looked like. It turns out the whole thing was made by AI. I was surprised, even though I still like the song.
I do use AI at school, through a program called SchoolAI. Mostly I put my writing in and it gives me ideas or helps me revise. You can’t have it just write for you, but you can use it to help. When I’m older I want to be an artist, or maybe a librarian. I’d probably use some technology either way. But the drawing and the songwriting? Those I want to keep doing myself.
The Cringe-o-meter (continued)
We asked kids to rate a range of AI uses from totally fine to not okay.
PING ZHU
The Pre-Premed
Evelyn, 13, Oregon
In January, I was diagnosed with type 1 diabetes, and that’s when AI became a bigger part of my life. Now when we’re cooking, we can run a recipe through ChatGPT, tell it the serving size, and it works out how many carbs there are. We use AI like that a lot.
My glucose monitor and my insulin pump also talk to each other using their own kind of AI to predict dosing. The monitor tracks what my blood sugar actually is, and the pump does the math. So if it predicts that my blood sugar will be high in 30 minutes, it gives me a correction dose, and if it predicts I’m about to go low, it stops the insulin before that happens. When I was first diagnosed I was still doing shots, and I went low almost every night. It was really stressful. Now the pump can catch it, and at night my phone goes off if I drop, so I wake up and drink juice. Mostly, I just get to sleep more because of it.
But the hardest part of having diabetes isn’t something I can use AI for. It’s remembering to carry all my supplies everywhere—to school, to a long day of anything.
I do use AI for school sometimes. Memory tricks when I’m studying for a test, ideas to get a project started. It’s a really good tool for that. But I don’t know exactly how I’ll use AI in the future. I want to be an endocrinologist someday, so I figure something will come up, since I’m already using it to help with my diabetes. I know other people worry that AI is going to take over the world. I don’t really think so. I still think we’re in control of it, and I think the benefits outweigh the risks.
I don’t know how exactly I’ll use AI in the future. I want to be an endocrinologist someday, so I figure something will come up, since I’m already using it to help with my diabetes.
The Inventor
Krishiv, 18, Ontario, Canada
When I was growing up, I always liked building things: Lego builds, Minecraft worlds, and then video games in Scratch. I’d make a little game, post it for other kids to play, read the comments, and make it better. Then, when I started high school, I had to spend way more time studying than I ever had, and honestly I just wanted to build things. So I went looking for ways to get good grades while studying less. Khan Academy had an AI tutor in the works, but it was stuck behind a waitlist, so I figured, why not build my own?
After months of launching random stuff, I created an AI tutor called Aceflow. You could feed it anything a teacher assigned—a 30-minute lecture video on YouTube, a blog post, a PDF of the textbook or presentation slides—and it would spin up endless practice questions, with a tutor on the side that explained things the way my teacher did. I built it just for myself, showed it to my friends, then put it on TikTok. It got tons of views on TikTok and thousands of users.
JUSTYNA STASIK
Was I worried people would call it cheating? Not really. I knew how to defend it: A tool like this isn’t so different from well-off families hiring expensive private tutors, except everyone gets one. That part mattered to me. Back in eighth grade, a teacher had me run a little computer science class for about 30 kids with special needs, and once they got personalized attention, they were building games nobody expected of them. That convinced me that kids are capable of so much more than people think, and AI can help scale that level of personalized attention to everyone. That unlocks so much potential.
That first AI tutoring project ended up helping me land part-time roles at BenchSci (one of Canada’s biggest AI companies) and Simple Ventures (a venture firm). More recently, I joined an AI lab at MIT; co-instructed an AI agents course with an MIT professor; and launched CheetahPrep.com, an SAT prep platform that uses AI to adapt to each student.
I’m generally optimistic about how AI will impact humanity, but when other kids’ first reaction is fear, I think that’s an important sign too. It’s a reminder that we should be excited about the future while still being mindful of the risks, working together to make AI work for humanity.
Correction (August 13): An earlier version of this article misstated Krishiv’s age. He was 18 at the time of publication.
Jen Swetzoff and Keeley McNamara are the founding editors of Anyway, an independent print magazine for tweens and teens.
Jos Benschop is climbing a ladder to get to the top of his newest machine.
It’s a bit of a schlep. The contraption is the size of a double-decker bus—more than 150 tons of gleaming precision-milled aluminum covered in thousands of snaking tubes, colored cables, and pressurized tanks. From the ground, it looks like a futuristic V8 engine. When I reach the top with Benschop we’re looking down from about 15 feet in the air, with bunny-suited technicians scurrying around below.
It’s more than 200 cubic meters of tech—“mechatronic devices that hold a few mirrors in a position with atomic precision,” he says, gesturing at the gargantuan apparatus. Benschop, a tall and grizzled 66-year-old, has spent over a decade working with his engineers to design this thing, but even so, he’ll sometimes look at it and go: Oh my God.
Benschop is the executive vice president of technology for ASML, a Dutch company that is the linchpin of the microchip industry. If you want to make powerful chips to power phones or AI, a lithography machine like the one we’re standing on is what you need to create increasingly tiny circuitry. Lithography is the art and science of shining light on a silicon wafer to pattern out the transistors, wiring, and other components of the microchips that will be cut from it.
The chipmaking field is essentially controlled by only two big players: ASML, which creates the lithography machines, and TSMC, the chipmaking giant.
Nine years ago, ASML began selling machines that use a daring new way of patterning chip features. These machines employ extreme-ultraviolet light, or EUV—radiation well outside the visible spectrum that they produce by shooting lasers at tiny molten drops of tin, tens of thousands of times a second. Those first machines—the result of an R&D moonshot that lasted 16 years and cost about $10 billion—can craft transistor features with a resolution of 13 nanometers. This new machine can do even better: It has a resolution of just eight nanometers, the width of about 40 silicon atoms. The devices are now shipping to chipmaking factories, or fabs, at an eye-watering price: $400 million each.
But chipmakers will fork that cash over, because they are in a desperate race to produce new and improved chips every year. That means getting their mitts on machines that can make ever smaller components and cram them together ever more densely—part of a long-standing recipe for creating faster and more energy-efficient chips.
For years now, ASML’s tools have been critical to keeping Moore’s Law alive. Without the company’s advanced chipmaking technology it is very possible that chip density—and the ability to perform ever more calculations—would have plateaued.
The AI industry has produced new and ravenous demand for denser chips, as firms like OpenAI and Anthropic scramble to erect server farms that train and deploy new, ever-more-powerful models, which require new, ever-more-powerful hardware. ASML’s latest machine promises to help keep the AI party raging for at least another decade.
“We can allow customers to go to smaller and smaller features, and that opens up the space for whatever we see now today in AI, which is absolutely mind-blowing,” Marco Pieters, ASML’s CTO, told me. “I think we’ve only seen the tip of the iceberg.”
Its relentless push for “shrink”—as they call it in the chipmaking industry—has made ASML a dominant force: The company produces about 90% of all chip-lithography tools worldwide. If you make chips, ASML is unavoidable.
But that monopoly position makes some people, and governments, uneasy. The chipmaking field is essentially controlled by only two big players: ASML, which creates the lithography machines, and TSMC, the chipmaking giant in Taiwan, which uses ASML’s machines to craft the vast majority of all microchips. This duopoly is so powerful that it has geopolitical implications. In an effort to prevent China from developing advanced AI, the US government pressured the Dutch government to impose an embargo in 2019: ASML isn’t allowed to sell high-end machines to any Chinese firm. Geopolitically, “chips are the new oil,” says Marc Hijink, the author of Focus: The ASML Way. Being deprived of them can be as disastrous as being deprived of oil. And in that metaphor, you might say, ASML is the Strait of Hormuz.
James Proud, the cofounder and CEO of the lithography startup Substrate, says the situation is not ideal. The US is “dangerously reliant” on a supply chain that’s overseas and increasingly pricey, Substrate says on its website. “There’s a huge concentration in a small number of players,” Proud says. “And the supply chain is just very expensive.”
Which is why, after two decades of ASML’s dominance, would-be competitors are now gunning for its territory. China is hungrily pouring billions into trying to replicate ASML’s tech. And startups like Substrate are trying to get in the game as well, setting their sights on creating lithography machines that are cheaper, smaller, and even more capable than ASML’s behemoths. Will any of them succeed? The near future clearly belongs to ASML, but as its engineers well know, you can unseat a giant with the right trick of the light.
Making chips is, oddly, a bit like silk-screening a T-shirt. To print a pattern on a silicon wafer, you start with a pattern on a reticle—a mask that carries the design. Shining a light on the reticle transfers that pattern to the wafer. The light interacts with a layer of chemicals on the wafer, fixing the pattern in place.
The size of a chip’s features is partly set by the wavelength of light the machine uses: The smaller the wavelength, the teensier the circuitry you can create. You can stretch the capabilities of a wavelength somewhat; increasing what’s known as the numerical aperture, which usually means swapping in a bigger lens, can further focus the light and thus lay down patterns for smaller and smaller components. Eventually, though, this trick hits its limit, and you need to find a new form of light with a smaller wavelength.
So the history of chipmaking has been a two-step dance. The industry finds a good source of light, eventually increases the numerical aperture, and then finally accepts the need for a smaller wavelength, starting the two-step all over again. Up to the early 1990s, chipmakers used visible light, with a wavelength of about 400 nanometers. By the mid-’90s they’d upgraded to deep ultraviolet, ultimately getting it down to a 193-nanometer wavelength. By the late ’90s they saw the end of the line approaching for deep ultraviolet. But what would come next?
All the options were troublesome. They could shift to x-rays, with a teensy one-nanometer wavelength, but they were devilishly hard to focus. Beams of electrons and ions were equally precise; but they worked like dot-matrix printers, transferring a pattern point by point, which was far too slow. (The chip industry wants a machine to crank out hundreds of wafers per hour.)
“It’s a very engineering-heavy company: Let’s send thousands of engineers and just have them mow down these problems. That’s what they did, and it worked.”
Jeff Koch, analyst, SemiAnalysis
Around 2001, ASML, then a smaller player in the lithography world, placed its bet on another option: EUV, with a wavelength just shy of the x-ray range. Nikon and Canon were working on it as well, but they dropped out—while ASML kept going. The idea was full of unknowns. Nobody knew how to reliably generate that type of light, nor how to focus it; EUV is absorbed by regular glass lenses. It’s even absorbed by air. ASML figured it would take six full years to wade through this R&D nightmare.
In reality it took those 16 years and about $10 billion in research, but it worked. The machine, which works in a vacuum, creates EUV light by vaporizing molten tin and using mirrors to direct it. Zeiss, a historic German optics company, had to invent new techniques for polishing and inspecting the mirrors, using an ion beam to knock off minute imperfections.
“They sort of ignored the buzz of, like, Hey, this is never gonna work, and they just beat their heads against these huge engineering problems,” says Jeff Koch, who used to work for ASML and is now an analyst for the chip-industry research firm SemiAnalysis. “It’s a very engineering-heavy company: Let’s send thousands of engineers and just have them mow down these problems. That’s what they did, and it worked.”
When the first EUV machines went on the market in 2017, they cost well over $100 million apiece. Some observers wondered whether the demand would really be there from the major chipmaking firms—TSMC, Samsung, and Intel. In the years chipmakers were waiting for EUV to happen, the lithography industry had developed clever ways to improve on old-fashioned deep ultraviolet light. (If you put a layer of water on top of the wafer, for example, the light could focus more narrowly.) Maybe EUV wouldn’t be much needed for a while?
But ASML lucked out. Only a few years after EUV debuted, OpenAI released GPT-3 and then ChatGPT. Artificial intelligence burst into the mainstream. Instantly, firms like OpenAI, Google, Meta, and Anthropic were hungry for increasingly high-end chips as they built massive server farms to train and deploy large language models. EUV made it easier and faster to crank out AI-tailored chip designs. Nvidia began producing elite GPUs—processors perfectly suited for AI training—that cost $40,000 a pop; the big companies couldn’t get enough. The AI wars were on, and EUV was in demand. In 2025, ASML says, it sold nearly 50 EUV machines to companies and pulled in nearly $40 billion in revenue. As of press time, the company’s market cap was over half a trillion dollars.
ASML’s new machines have no shortage of potential customers. But there is one in particular, with deep pockets, that can’t buy them for any amount of money: China.
The US wants to hobble China’s ability to create cutting-edge AI chips—or any advanced chips, for that matter. So when ASML began selling its original EUV machines, in 2017, the Trump administration successfully pressured the Dutch government to forbid the company from selling them to any Chinese firms. The US had also imposed export controls on China’s telecom giant Huawei, banning US firms from using its 4G and 5G equipment.
This one-two punch incensed the Chinese government and stirred it to action. China is now pouring billions into catching up and trying to develop its own EUV chip-patterning technology. A Reuters report last winter found that a government skunkworks employing former ASML staffers had cobbled together a machine so huge it filled the entire floor of a lab. It’s unclear how well it works. The experiment may well be making some chips, says Hijink, but he doubts it can do so at an industrial scale.
A mirror is installed in an optical system for the high-NA machine.
COURTESY OF ZEISS
Officially, the government denied it was pushing to develop EUV tech. An editorial in the Global Times—a newspaper closely allied with the Chinese government—pooh-poohed the report, claiming that China was still happy to work with the West to get access to chips. “Our goal has never been to build a self-sufficient ‘technology island’ in isolation,” it stated, “but rather, on the basis of achieving autonomy and control over key technologies, to integrate more deeply and equally into the global innovation network.”
Experts say the reality is in the middle. China definitely craves a domestic ability to make high-end chips. And unlike ASML, it doesn’t need its EUV machinery to be efficient and profitable, cranking out about 200 wafers an hour. Any output would help wean it off reliance on the West.
“They would be very happy to have a tool that does one wafer per hour and it costs them a fortune to run,” Koch says. “They would build a fab with a thousand of those and be super happy with it.”
Still, producing and managing EUV light well is a feat that might take years, some told me. In the meantime, the Chinese will lean hard on deep-ultraviolet lithography, developed in the ’90s, making the most of an alternative but slower approach known as multi-patterning, says David Lin, senior advisor for tech leadership at the Special Competitive Studies Project, a think tank that focuses on security and technology. “They’re going to push DUV to the absolute limits,” Lin says.
The AI race is also pushing China to devise ever cleverer ways of developing LLMs that don’t rely on the fastest AI chips. In the US, OpenAI, Anthropic, and Google are fighting over who can buy the biggest piles of hot Nvidia chips. Since China can’t compete that way, it is innovating not in hardware but in software—building lighter-weight LLMs like DeepSeek.
As China rumbles into action, ASML has remained laser focused on shrink. To go even smaller, Benschop and his engineers decided, they wouldn’t shift to a new form of light. They’d do the second part of the two-step: They’d raise the numerical aperture of the machine by more than half (for those keeping track of the specific numbers, it would be a switch from an NA of 0.33 to an NA of 0.55). That would let them cut the size of the transistors by close to half and nearly triple their density on a chip.
This would also be an easier climb. Without the need to develop an entirely new source of light, the new machine—based on high-numerical-aperture EUV, or “high NA”—would be evolutionary, not revolutionary.
Still, building the new system did present a few gnarly challenges. In an EUV machine, the way you transfer an image onto a wafer is by shining light at the microchip pattern on the reticle and then using an optical system to take the reflected light and demagnify that pattern, shrinking it down to the size you want on the wafer. The light hits only part of the reticle at any given time, so you quickly move the reticle back and forth to expose every part of the pattern to the light.
Going to a higher numerical aperture meant they could have smaller features on the reticle. But this also meant that some of the light would be arriving at the reticle—and reflecting off it—at a steeper angle.
That’s what caused problems. The pattern on the reticle is three-dimensional, so light arriving at such a steep angle caused shadows—much the way slanted sunlight creates shadows in the Grand Canyon. That stood to diminish the machine’s ability to make clear patterns.
The new reticle moves with acceleration up to 22 g, much faster than in the company’s original EUV machine. “Don’t try to sit on it, because you’ll pass out.”
The solution was to change the pattern on the reticle—along with the way the mirrors took the light and shrank it down to impart the pattern to the wafer. The designs on the reticle would now be twice as long as they were wide—stretched, as it were, in one dimension.
But this design came with its own problems. The changes to the mirrors meant the area on the wafer exposed during a single scan was half the size it was with the original EUV machines, reducing the system’s speed. And ASML couldn’t tolerate any slowdown: Chipmakers were paying it for machines with massive throughput, about 200 wafers an hour.
If one part of the system slowed down, another part would have to speed up. The engineers decided the machine should move the reticle faster, which meant making the entire mechanism lighter and dramatically redesigning it. The new reticle moves with acceleration up to 22 g, much faster than in the company’s original EUV machine. “Don’t try to sit on it, because you’ll pass out,” Pieters told me. The wafer stage moves around faster as well, in tandem with the reticle.
Meanwhile, over in Germany, Zeiss’s engineers were busy designing mirrors to accommodate the higher numerical aperture and asymmetric shaping of the light. The new mirrors would be about twice as large as those in the regular EUV machines, and the projection system, which carries light from the reticle to the wafer, weighed fully 12 tons, seven times more than before. Zeiss built a new robot-assisted production line to handle these ponderous new beasts. The company says they’re the smoothest surfaces they’ve ever made.
At the same time, ASML was working on making its EUV light source even more powerful, to help make the wafer-exposing process go faster. The engineers calculated that they could improve the output of EUV if they hit each tin droplet three times with the laser instead of twice, as they do in the first machine. That meant the already-hectic system of firing tin would need to speed up by 50%. “The lasers just keep getting bigger,” says Alex Schafgans, the head of engineering at ASML in San Diego, where the EUV light source is built.
Indeed, the lasers for a single machine now fill an entire room. After Benschop showed me the massive high-NA device, we walked across the hall and entered a chamber filled with hulking six-foot-tall boxes that were part of the laser system. Peering through tiny windows in the sides of the units, we could see the glowing purple plasma used in creating the laser light.
When high-NA machines began to roll off the assembly line, one company was waiting hungrily: Intel. The company purchased the very first high-NA machine put up for sale, and in the spring of 2024, 300 ASML engineers showed up in Oregon at one of Intel’s fabs to begin assembling and testing it.
“ASML actually put a giant ribbon around one of the boxes,” says Mark Phillips, an Intel fellow who is director of its hardware and lithography solutions, laughing. His team has been testing the machine to see how well it performs; Phillips wouldn’t give details other than to say he’s “very pleased at the rapid pace of tool health.” He also wouldn’t give a date for when Intel would start using it to make chips, though observers say that will likely happen next year. The company plans to ease it in, using it for just a few precision components on a chip and then gradually for more and more.
What’s at stake is a chance to recapture its mojo. Intel was once a silicon powerhouse, designing the most cutting-edge CPUs for computers and servers, and building them in its own fabs. But in the 2010s, the big new markets were mobile-phone chips and GPUs for AI and gaming, and Intel rapidly lost ground. Apple designed its own mobile chips (and had TSMC make them), while Nvidia did the same thing with GPUs. Google began banging out its own TSMC-made AI chips called TPUs in 2015, and soon it was stuffing data centers full of them.
Intel fellow Mark Phillips briefs members of the media on the high-NA tool at the company’s Fab D1X in Hillsboro, Oregon. Intel was ASML’s first customer for the new EUV machine.
COURTESY OF INTEL CORPORATION
So in 2021 Intel announced a moonshot. It would aggressively begin building out a foundry business, one that would go toe to toe with TSMC. Instead of creating Intel chips, the Intel foundry would manufacture designs for customers like makers of mobile phones and AI chips.
Intel hopes that being the first to wield high-NA technology will give it an edge in the silicon rat race, making it possible to print tiny patterns faster than anyone else.
It could also make things simpler for customers. Over the years, while waiting for EUV machines to emerge, chip designers used multi-patterning to squeeze more life out of the older forms of light. Every chip is made out of layers, which are laid down to make components like the switches and wiring. If you’re working on one of those layers and need to make features tinier than your machine can normally produce, you can break the pattern for that layer up into several patterns and then expose the wafer to them one at a time. This strategy helped chipmakers keep using older (and cheaper) machines while still creating tinier and tinier components. But multi-patterning is a hassle: It’s more challenging to design the complex overlay of patterns, and much slower to print each chip. Designing a chip is far easier if you know you can do “single patterning,” blasting each layer in one go.
Observers say it won’t be easy to build a foundry business that bests TSMC and Samsung on their own terrain. “Leapfrogging is difficult,” Hijink says. But it’s also true that the high-tech world has such a ravening hunger for better chips that Intel could succeed, simply because even TSMC and Samsung can’t fulfill all that need.
“There’s spillover demand, so Intel can survive off that,” Koch says. “It’s not even scraps now. It’s a meal. It may not be the best foundry, but they can make chips, and there’s only three companies that can do that, right?”
TSMC, for its part, seems to be biding its time when it comes to high NA. “TSMC will deploy high-NA EUV when it is mature and ready to deliver maximum benefit to our customers,” the company wrote to MIT Technology Review. Some suspect it won’t use the machines in serious volume until the 2030s. Part of the reason is cost: TSMC is ruthlessly focused on producing chips as cost-effectively as possible, and the high-NA tools are a blistering $400 million each, far more than the previous EUV rigs. And unlike those, the new machines are not a revolutionary leap upward.
“This is like 30% to 50% better in terms of capability,” says Koch, the analyst and former ASML employee. “This is probably the first tool that hasn’t obviously made business sense right away for ASML.”
It’s not that the industry won’t eventually embrace high NA en masse, Koch says. Most companies will need to, if they want to keep going smaller. But TSMC is more likely to push ahead as far as it can go with its existing EUV tools, using onerous multi-patterning to wring as much as it can out of that generation until it absolutely needs to switch.
“The industry has only shifted paradigms when it just absolutely cannot extend—even one more little bit—out of what it’s been doing,” Koch says.
China isn’t the only party looking to upset the current balance of power. The dominance of ASML, and the swelling cost of its tools, is prompting other upstarts too. But instead of trying to replicate ASML’s breakthroughs in EUV, they’re doing an end run—working on lithography tools that use entirely different forms of light. These will be far cheaper, they promise, and just as powerful.
One is Substrate, a San Francisco–based startup. Founded four years ago, it’s working on a tool that uses x-ray light produced by a particle accelerator. X-rays have a remarkably tiny wavelength, making them a potentially powerful way to create minute features.
Particle accelerators have historically been enormous, making them difficult to fit into a chipmaking process. Substrate says it has harnessed decades of scientific improvements in particle acceleration to produce a light source that’s smaller and suitable for mass production.
Last year the company released images showing that it had created fine patterns, which Proud, the CEO, says are only possible now with a high-NA EUV machine. He says Substrate’s goal is to produce chips at scale by 2030.
But Proud doesn’t intend to sell the tools to TSMC or Intel. Indeed, he doesn’t plan to sell them to anyone. Instead, Substrate wants to create its own fab, building chips using its own tools.
“The amount of chips we’re going to need is going to be many orders of magnitude larger than even the wildest projections you have now.”
James Proud, cofounder and CEO, Substrate
The semiconductor industry, Proud argues, needs new approaches, because it’s become too pricey and too centralized. A single fab today can cost $25 billion to build, up from about $5 billion in the 2010s, the company notes. It’s driving the cost of a single wafer full of advanced chips up toward $100,000, Proud says.
“That is, I think, a prohibitive cost,” he says. There also isn’t enough capacity in the supply chain: “It’s relatively slow and hard to flex to the current increase in demands.” He admires ASML’s EUV tooling—it’s “the apex implementation of that technology”—but new approaches are needed.
That’s partly for national security reasons. Proud and his team think it’s too dangerous for the US to rely on foreign supplies. But he also predicts the current AI boom will go into overdrive, creating a massive demand for chips that the existing ASML/TSMC duopoly won’t be able to deliver: “The amount of chips we’re going to need is going to be many orders of magnitude larger than even the wildest projections you have now.”
ASML’s machines use lasers and molten tin to generate the EUV light.
CHRISTOPHER PAYNE
Substrate predicts it will be able to produce finished wafers at $10,000 a pop—a tenth of where Proud predicts the rest of the industry is heading. Proud says that’s partly because the company’s system will be vertically integrated, so it will control all parts of the chipmaking process, but also because its lithography tooling will be less complex: “We’re able to put together in a sort of simpler package.”
Still, Substrate is playing its cards close to its chest. Unlike ASML, the company isn’t offering nuanced detail on how it generates light, or on how that then translates into making patterns on a wafer.
Substrate’s ambitions give some industry observers pause. Hijink, who thinks it is probably “unachievable and impossible” to simultaneously master both a new form of lithography and high-throughput fab techniques, regards the company’s secrecy as a red flag. “This industry is about open innovation,” he says.
Koch is more impressed by its ambitions and funding. The type of technology it’s pursuing “is really cool,” he says. “It’s interesting.” But “there’s a long road between lab-scale demonstration and high volume,” he adds. “Is this like an imminent disruption to ASML? Probably not.”
Another startup that is aiming to hit the market around the same time as Substrate is Lace Lithography. Based in Norway, it is devising an entirely different approach—one that doesn’t use light at all. Instead, an energized beam of helium atoms is pointed at the pattern on the reticle. When the helium atoms then hit the wafer, the atoms transfer their energy to it, imparting the design to the chip.
The idea dates back a while. Bodil Holst, the CEO, took it up in 2008, when she was a physicist studying the use of atom beams. MIT professor Henry “Hank” Smith, a pioneer in using x-rays for lithography, told her she should explore using atoms as a mechanism for making microchips, because back then he wasn’t sure ASML’s EUV moonshot would work. “Even if it does, we’ll need atoms eventually,” he told her.
Holst did some experiments to investigate the idea further and partnered with a former PhD student—Adrià Salvador Palau, a physicist and expert in machine learning—to found Lace. Like Substrate’s, its tool is completely different from ASML’s massive machinery. The source of the excited atoms “looks a bit like a rocket motor,” says Palau. “It’s very cool.” While EUV’s wavelength is 13.5 nanometers, the helium atoms offer a precision of 0.1 nanometers. The process also requires far less power, and the machine is intended to be far smaller. Holst tells me the company aims to have machines ready to sell to fabs by 2029 or 2030.
“I think everybody’s really looking forward to something that extends a road map beyond light, beyond EUV,” Palau says.
ASML is watching these upstarts with curiosity. Benschop says he can’t assess whether Substrate’s technology will work reliably and affordably, because the company hasn’t explained anything about its processes. But he went to a conference where Holst and Palau did a presentation outlining Lace Lithography’s technology.
“I’m incredibly impressed with how they do it,” he says. The problem, he says, is he doesn’t think the process produces patterns on the wafer that are deep enough to be useful. “I cannot see how they would scale it to a viable volume product,” he told me.
He suspects ASML’s mastery of EUV will keep it on top for the near future. “So far, I have not seen a viable alternative,” he says. He thinks there’s “no serious runner-up” when it comes to volume manufacturing of the most advanced chip generations.
It’s true that major shifts in chipmaking are slow, says Chris Miller, a professor of international history at Tufts University and the author of Chip War, a book about the worldwide struggle for dominance in the industry. “No doubt we’ll eventually have alternatives [to EUV],” he told me via e-mail. “But it’s worth noting that lithography transitions have historically taken years, if not decades.”
ASML’s executives, too, are pondering their future. Benschop expects high-NA technology to dominate chipmaking into the 2030s. Beyond that? The industry has, indeed, tended to shift to a new form of light every decade.
“You may argue it’s time for the next decade,” he told me after we’d stripped off our bunny suits and he was relaxing with a coffee.
But ASML’s executives suspect they can continue to squeeze more capabilities out of EUV by increasing the numerical aperture even further on their existing machine. They’re already toying with a design that would take an NA of 0.55 to an NA of 0.75: “hyper NA.” It could let them pattern wafers with a resolution of six nanometers. They’re also working on standardizing their various optics into a platform of a single size, so customers could order one machine outfitted for either regular EUV, high NA, or hyper NA. If it’s all in the same-sized unit, it would simplify the costs and logistics of integrating each into a fab. If the company goes through with it, Benschop figures, the hyper-NA tool might hit the market seven or eight years from now and be sold in volume during the second half of the 2030s.
For now, the ball is in ASML’s court. “We’re pushing the limits of physics,” Pieters told me. The question now is whether anyone else can push harder.
At the end of a tense and scoreless first half of a soccer match between the English men’s team and rival Germany, millions of Brits let out a collective sigh and did what they so often do in moments of stress: They made tea. That wave of electric kettles clicking on, however, caused a different kind of stress: a huge and sudden increase in demand for electricity. But National Grid, which operates the local transmission network, was ready.
Just as those kettles started heating up, an AI program sent instructions to a data center in London to slow down some of the facility’s power-hungry chips. This reduction helped make sure there was enough supply to match demand, staving off potential blackouts or damage to electrical hardware. For data centers, which normally guzzle power without consideration for anyone or anything else’s needs, it was a radical departure.
It was also a test. The software was controlling a real data center, but there was no game happening at the time, in December 2025. Engineers were doing a trial run for a new breed of data center built to be flexible about its electricity needs, so they re-created the energy demand facing the UK’s grid during a match from the 2020 Euro tournament. They wanted to see how their software, called Conductor, would have responded had it been online at the time.
Conductor is the signature product of Emerald AI, a firm based in Washington, DC, that’s part of a wave of companies trying to figure out whether data centers can work within the confines of the existing electric grid.
This year, Emerald is set to deploy Conductor in a new facility in the part of Virginia known as Data Center Alley, this time connected to the live grid. When overall demand spikes, Conductor will turn down the power used by the data center, while making sure its servers still carry out their timeliest and most important jobs. Emerald’s partners on the project—which include Nvidia and the giant data-center operator Digital Realty—bill it as one of the world’s first “power-flexible AI factories.”
Demonstrating that data centers can participate in this kind of give-and-take could ease what many tech leaders identify as the bottleneck in getting facilities online: It takes far longer to get approval for, construct, and connect new power plants than to build data centers. PJM, the grid operator in Virginia and the largest one in the US, for instance, needs eight years to bring new generation online, according to RMI, an energy research and advocacy group. “We need to solve the energy equation,” says Josh Parker, head of sustainability at Nvidia. “AI factory flexibility is the bridge between the incredible demand for AI and the immediate limitations of our energy grid.”
Speed, though, is only one of the issues. Once facilities do plug in, neighbors often criticize them for drawing too much electricity and contributing to rising prices. They say the data centers generate more noise than they do long-term jobs, contribute to pollution, and threaten to put people out of work. Organizers stalled over $150 billion worth of projects in 2025, according to Data Center Watch, and policymakers alert to the public mood are starting to impose limitations on development.
More than a dozen states are considering bans, and local moratoriums are in effect in places like Minneapolis and DeKalb County in Georgia. At the federal level, the GRID Act, a bipartisan bill in the US Senate, proposes to sever new data centers from public grids entirely. Some operators are already moving that way by trying to develop their own power generation.
Rather than rushing to build new power plants, companies could find part of the solution to the crunch right under our noses—or, more precisely, in the transmission lines under our feet and above our heads. The existing system operates near its full capacity during only a small number of high-demand hours throughout the year. This means, some grid experts argue, that if data centers can limit the power they draw during those stretches, they won’t need to wait for big infrastructure upgrades or build their own off-grid generation.
Indeed, a growing number of studies have shown there could be plenty of power available for data centers that can flex. A widely discussed 2025 report from researchers at Duke University found that the US grid could offer an additional 76 gigawatts—about 5% of its entire capacity, and about enough to accommodate projected data-center growth in the US through 2030—to facilities that are willing to reduce their usage just 0.25% of the time. That’s about 22 hours a year. And when researchers from Princeton University and two grid-modernization companies looked at locations for new data centers in the PJM region, their report, which was funded by Google, found that a 500-megawatt facility capable of flexing for less than 1% of the year could reach full operation three to five years faster than one that’s inflexible.
Flexible power connections could also help data centers address some of their PR problems. By decreasing their draw at times of grid stress, for instance, they could avoid diverting power from where it’s most needed, thus boosting stability. By using existing capacity, they might be able to reduce the need for new fossil-fuel power plants and spread fixed costs over more electricity users, pushing prices down.
The AI power pinch is attracting resources and research into strategies for grid flexibility overall, which could help negotiate a tricky period: Taken together with electric vehicles, air-conditioning, and other sectors, data centers are helping drive what analysts predict will be a 25% increase in US electricity demand by 2030 compared with 2023 levels.
Ideally, flexibility gives grid operators more control over the flow of electrons, making them leaders of a harmonious ensemble rather than hostages to inflexible electricity requirements. That will help them manage demand spikes across the entire system and deal more effectively with the intermittent nature of renewables like wind and solar. “Demand flexibility is incredibly useful for power grids,” says Johanna Mathieu, a grid expert at the University of Michigan. “It helps reduce electricity costs and improve grid reliability.”
But while advocates see plenty of benefits, the concept brings complexity. For data centers, compromising on energy needs can be a hard sell. Flexibility requires utilities and grid operators, which tend to be operationally conservative, to change long-held practices. And some skeptics also say that flexibility distracts from the very real need to build more grid infrastructure faster, and could even pose risks to our electricity supply.
Still, some technologists, grid operators, and utilities are hoping to show that flexibility works—not only in white papers or simulations but in real life.
The poster children for data-center growth default toward inflexibility. Hyperscalers like Microsoft and Oracle have proposed enormous new centers, many of which would rely on off-grid, natural-gas-burning power plants. When xAI wanted to speed up the buildout of the Colossus site outside Memphis, Tennessee, it rolled up with gas turbines on flatbed trucks. The facility, now in operation, is facing blowback from regulators and residents about the spike it’s causing in emissions and other pollution. In any case, there aren’t enough gas turbines worldwide to meet the demand from data-center operators.
One big obstacle for anyone demanding a lot of power is that our grids are mostly rigid. They’re designed to supply enough power to meet total demand when it’s highest, even if that’s for only a relatively small number of hours a year. That conservative approach is a simple route to reliability, but it means that the grid has quite a bit of headroom. “The grid is already overbuilt by a lot. If you were an airline running at 30% utilization, you would not buy more planes,” says Amit Narayan, the cofounder and CEO of GridCare, a company developing flexibility technologies, referring to a 2025 Stanford study of transmission lines in western North America. “If you are running a grid at 30% utilization, there’s no scientific reason you can’t go to 60.”
“If you were an airline running at 30% utilization, you would not buy more planes. If you are running a grid at 30% utilization, there’s no scientific reason you can’t go to 60.”
To be fair, the idea of flexibility isn’t entirely foreign to grid operators. For decades, they’ve practiced a technique called demand response: When it looks as if demand will get too close to supply, as it might during a heat wave when many people turn on the AC at the same time, they call large commercial or industrial facilities and ask them to shut down parts of their operations. This method can help avoid the need to fire up so-called peaker plants, which run on fossil fuels, but it’s slow, imprecise, and hard to scale.
In the 2000s, as the adoption of technologies like electric cars and solar panels presented new challenges, more internet-connected grids also provided new means of flexibility. Virtual power plants, or VPPs, offered a smarter, faster, more granular alternative. Electricity customers ranging from factories to homeowners with smart thermostats, solar panels, or big batteries would allow the utility to adjust their draw to help meet demand—often getting paid for their (frequently unnoticed) trouble.
After the generative AI boom began with the release of ChatGPT in 2022, some companies began to see flexibility as a way to get data centers set up more easily, efficiently, and affordably. If they bring AI money into existing grids and reduce or defer the need for expensive upgrades, data centers could actually help spread out fixed costs so as to lower rates for other users. A study from Duke University published this past February, for instance, found that flexibility could reduce rates by 0.5% to 2.8%.
PETRA PÉTERFFY
The trick is figuring out how data centers, notorious power hogs, can keep operating when their flexible connections are throttled. Flexibility specialists envision three possible ways. The simplest is for the new data center to install on-site backup power storage or generation to tap when the grid is maxed out—at their own expense, of course.
A facility could also fill the gap by drawing on a VPP. The utility would turn down the electricity going to users who signed up for the VPP, and the data center would pay them for their flexibility. This method wouldn’t require any major infrastructure, but it would require the utility to have a big VPP program and to coordinate the exchange at a time when the grid was under stress. While VPPs exist to some extent in nearly 40 states, the rules governing them vary widely, and they are empowered to do more in some areas than in others.
Finally, a data center could simply use less power at peak times. The conventional wisdom is that they won’t go for such limits, particularly when every number-crunching server can feel like a goose potentially laying little golden eggs. But some experts are betting that the value of getting up and running quickly is enough to change their minds. “There is a clear and growing trend,” says Ayse Coskun, chief scientist at Emerald AI. “Operators are increasingly willing to trade some level of flexibility for faster grid interconnection.”
GridCare, a startup based in Silicon Valley, was one of the first companies to use flexibility to get data centers online quickly. Instead of looking at grids only in worst-case scenarios when electricity demand is highest, the company analyzes the system under all conditions, explains CEO Narayan, who studied smart grids at Stanford. It feeds every part of the grid—including power plants, lines, substations, and homes—into a generative AI model that creates a “digital twin” for different grid configurations. It then picks out results that could unlock capacity while maintaining reliability, and it feeds those into another model trained on the physics of electrical components like resistors and capacitors to make sure they’re realistic.
GridCare found its first customer in the Silicon Forest, an area in the Pacific Northwest named for the trees that dominate the landscape and the IT industry that has more recently sprouted up there. The local grid needed more capacity to support more data centers. “Data centers wanted ‘speed to power,’” says Isaac Barrow, a manager of data-center relations at Portland General Electric, or PGE, the local power generator and distributor, “but transmission buildout is a long process that’s very costly.”
In 2024, Aligned Data Centers came to PGE wanting to expand its operation in Hillsboro, Oregon, and PGE followed a recommendation from GridCare. Aligned will install a 31-megawatt battery, set to be in service in May 2027, and decrease its draw by up to that amount when the grid becomes congested. Bundled with other flexibility measures, that battery has allowed PGE to increase the capacity it can offer Aligned and other nearby operators by 80 megawatts without any new power plants. Though the buildout of data centers in Hillsboro has faced plenty of pushback from locals, Barrow points out that it could have the knock-on effect of lowering costs for ratepayers, because it spreads out the tab.
Other companies are promoting different flavors of flexibility. Google has been moving processing loads from facilities in areas experiencing demand spikes to those in less stressed spots since 2023. It’s signed agreements with five utilities, including the Tennessee Valley Authority and Indiana Michigan Power, that add as much as a gigawatt of flexibility.
Voltus, a major VPP provider across the US and Canada, markets a “bring your own capacity” program in which a data-center company can fund a VPP nearby. The grid operator can use the VPP to decrease demand at busy times, and participants get a financial thank-you. “We can spin up new VPPs on the order of months,” says Emily Orvis, Voltus’s vice president of energy markets. In June, the company signed their first such data-center deal: a three-year plan in which Google will bankroll a VPP in the PJM interconnection.
Of all the approaches to flexibility, Emerald AI’s may be the most ambitious: asking data centers to dial into the grid’s needs. The company’s Conductor software, which can run on premises or in the cloud, builds on the research of chief scientist Coskun. Her group at Boston University showed in a pair of 2013 papers that a data center could watch the grid and help balance big power fluctuations, such as the intermittent effects of solar and wind power. By 2022, she and her colleagues had tested their methods on a cluster of 36 research servers and shown that the system could respect power limits without breaking the processes it was running.
One of the most important questions for Conductor is deciding which AI processes can be slowed down to save energy without kneecapping performance. A lot of companies label their jobs by priority—a real-time chatbot query, for instance, might outrank something like a web search that’s part of a deep research project. When they don’t, Emerald AI tries to infer priority from the nature of the job. Conductor then analyzes the AI workload to determine how tweaking the power to a given processor will affect the performance and help meet the usage limits set by the grid operator.
“The performance curve changes for different kinds of workloads,” says Coskun. “Each AI job is going to have a different location on that curve. Our intelligence is figuring out where you are on that curve.”
PETRA PÉTERFFY
Last year, Emerald AI began assessing the technology’s readiness for real-world use in a series of tests, raising the difficulty each time. The trials were carried out in partnership with the Data Center Flexible Load Initiative—a collaboration among tech companies like Google and Nvidia, utilities like Duke Energy, and grid operators like PJM that aims to help establish a repeatable framework for power-flexible data centers.
The first challenge was in Phoenix, a fast-growing computing hub. For the test, Conductor took control of a group of server racks laden with 256 Nvidia A100 GPUs—hardware that can use about as much power as around 170 US homes. When presented with a simulation of a busy grid, Conductor reduced the power to the chips by 25% for three hours, while maintaining acceptable computing performance. Emerald AI and its partners reported the results in a paper in Nature Energy in December 2025.
The next trial forced the system to juggle surprise grid fluctuations without advance warning and redirect AI jobs from a data center in Virginia to a less busy one in Chicago. Then, in London, Conductor took the reins of equipment beyond the main GPU processors and faced a more complicated mix of fluctuations, including very short and long bouts of congestion—plus the notorious teakettle effect.
The progress so far shows that flexibility can work, at least in some situations, but only a small fraction of operators have pursued it as yet. “We’re just in the beginning innings of the game,” says Jesse Jenkins, one of the authors of the 2025 Princeton study and cofounder of Firma, a startup that works on data-center flexibility. “People are recognizing that this is a potential solution. The motivation is there; there are some bespoke examples. But there’s no uniform solution set that’s the default option, which is where we need to get.”
While data centers are going up across the US, no place on Earth comes close to the accumulated computing muscle in Northern Virginia’s Data Center Alley. The region is home to around 500 compute-crunching facilities, which represent 13% of the entire world’s capacity; the next two hot spots, Beijing and Oregon, contain 6% each.
There are proposals to build hundreds more facilities in Virginia, but a government study found that the state’s electricity demand will increase 183% (around 26 gigawatts) by 2040 if they all go forward, and supporting even half would be difficult. The power-flexible data center that Emerald AI, Nvidia, Digital Realty, and their partners are building in the suburb of Manassas could demonstrate how data centers can squeeze the power they need out of existing capacity. The facility, slated to come online later this year, is intended to give Conductor the chance to manage power at the largest scale yet and to respond to conditions on a live grid for the first time. In the UK demonstration, Conductor managed a 130-kilowatt AI cluster; in Manassas, it will pull the strings of a 96-megawatt hyperscale AI factory.
Some degree of flex will play a key role as we transition away from fossil fuels and toward a future that has to juggle technologies like solar and wind power, batteries, and electric cars.
For PJM, the Manassas facility points to a potential path through the current power crunch. “We think data-center flexibility, in different forms, will be essential for the reliable integration of data-center load over the short to mid term,” says Scott Baker, who manages demand-side markets at PJM.
But not all grid experts are so sanguine. PJM’s market monitor, which oversees the grid operator, says there are no workarounds when it comes to adding capacity. “The notion that large amounts of data-center load can be added without adding new generation is magical thinking,” says Joseph Bowring, an economist and the head of PJM’s market monitor since 1999.
One problem, he says, is that there’s no way to guarantee that a data center will actually take less power when demand is high. That is, absent any legal or regulatory push for flexibility or compliance, the utility won’t be able to step in to help prevent, say, a blackout. Utilities can rely on resources like power plants, but they can’t control or rely on data centers. “They do not want to be fully interruptible,” Bowring says of the facilities.
Stephen Empedocles, an advisor for technology companies, views flexibility as more of a tool than a silver bullet. “These approaches are excellent for improving grid reliability and getting more out of the infrastructure we already have,” he says, “but they are optimization tools.” They’re not substitutes for the “generation, transmission, and distribution expansion that will still be required,” he continues.
Flexibility advocates agree that over the long term, whether or not AI continues to boom, electrification will drive a need for more generation and transmission. Some degree of flex will play a key role in using grid infrastructure better as we transition away from fossil fuels and toward a future that has to juggle technologies like solar and wind power, batteries, and electric cars. A report published by the International Renewable Energy Agency in January 2026 found that grids around the world will need three times as much flexibility in 2030 as they had in 2019—and 10 times as much by 2050—to balance increasing demand with fluctuating supplies of renewable energy.
The challenge of powering AI could provide just the spark we need to do the work of designing and building smarter, more flexible grids, says Coskun. “I think with a crisis like this, there’s no quick solution,” she says. “Sometimes a crisis like this creates an opportunity to do something differently.”
Amos Zeeberg is a freelance science and technology journalist based in Bucharest. He’s developing a book about technology networks, including electric grids.
This story was updated on June 20, 2026 to clarify details about Emerald AI’s test in London in 2025.