Reading view
How much of the AI data center boom will be powered by renewable energy?
CaPow partners with Manufacturing Revolution to deliver in-motion charging to U.S. Midwest
1 million business customers putting AI to work
Scale Biology Transformer Models with PyTorch and NVIDIA BioNeMo Recipes
Training models with billions or trillions of parameters demands advanced parallel computing. Researchers must decide how to combine parallelism strategies,...
Training models with billions or trillions of parameters demands advanced parallel computing. Researchers must decide how to combine parallelism strategies, select the most efficient accelerated libraries, and integrate low-precision formats such as FP8 and FP4—all without sacrificing speed or memory. There are accelerated frameworks that help, but adapting to these specific methodologies…
Inside Hyundai’s Massive Metaplant

When I traveled to Ellabell, Ga., in May to report on Hyundai Motor Group’s hyperefficient Metaplant—a US $12.6 billion boost to U.S.-based manufacturing of EVs and batteries—the company’s timing appeared solid. At this temple of leading-edge factory tech, Ioniq 5 and Ioniq 9 SUVs marched along surgically spotless assembly lines, giving the South Korean automaker a defensible bulwark against the Trump administration’s tariffs and onshoring fervor.
But dark clouds were already gathering. Consumer adoption of EVs had started slowing. The U.S. federal government’s $7,500 clean-car tax credit, which had helped hundreds of thousands of people make the leap to EVs, was being phased out.
Held securely on a yellow jig, a three-row Ioniq 9 SUV glides from station to station in the assembly hall. A view from below shows its generous, 110.3-kilowatt-hour battery pack, which, as in most EVs, sits below the floor of the car. The pack, which is shielded to prevent or limit damage in a collision, is part of an advanced 800-volt architecture for ultrafast DC charging. Christopher Payne/Esto
Near the Savannah-area factory, I drove a smartly designed Ioniq 9, a three-row SUV tailored to the United States’ plus-size tastes. I also saw a battery plant taking shape: a $4.3 billion joint venture between Hyundai and LG Energy Solution, on track to produce lithium-ion cells for Hyundai, Kia, and Genesis models in 2026. That facility is one of 11 low-roofed buildings that encompass 697,000 square meters (70 hectares), their pale green walls designed to blend into the Georgia countryside.
Backed by $2.1 billion in state subsidies, the Metaplant is the largest public development project in Georgia’s history. Covering 70 hectares, it is the centerpiece of Hyundai’s $12.6 billion total investment in the state, including the battery factory built with LG Energy Solution that ICE and other agents raided in September. Christopher Payne/Esto
That battery plant made headlines in September, when U.S. Immigration and Customs Enforcement (ICE) agents staged a workplace raid that led to more than 300 South Korean workers being detained and deported.
The episode highlighted the transnational cooperation—and tensions—inherent in importing a leading-edge manufacturing operation, a duality that might be familiar to anyone old enough to recall Japan’s game-changing entry into the U.S. automobile market in the 1970s and ’80s. The Metaplant is the largest publicly backed project in Georgia’s history. Its creation was accelerated by the Biden administration’s pro-EV policies, and it was also the centerpiece of Republican Gov. Brian Kemp’s bid to make his state “the electric mobility capital of the country.” Now, it was suddenly the latest flashpoint in an ongoing culture-and-trade war.
Automakers roll with the punches because they have no choice
An automated guided vehicle (AGV) prepares to pick up a rack of windshields from an automated trailer unloader, for “just in time” delivery to an assembly line where Ioniq 5 EVs are being built. There is no human intervention from the time parts arrive at the Metaplant’s loading docks to their installation. Christopher Payne/Esto
Robots perform myriad tasks, yet human hands are still best for precision work. Jerry Roach, the Metaplant’s assembly manager, says, “I want my people doing craftsmanship. I want to pay people well for the things humans do well, and take away the stuff that’s tedious and boring.” Christopher Payne/Esto
As with other EV makers facing hurricane-force headwinds, including the U.S. rollback of pollution and fuel-economy rules, Hyundai has chosen to forge ahead with its long-laid plans. Company executives call the Metaplant North America’s most automated car factory and the most advanced full-scale factory among Hyundai Motor Co.’s 12 global manufacturing facilities. It rivals or surpasses Japan’s most advanced plants, such as the best operated by Toyota. Compared with the near-Dickensian Detroit auto factory that I toiled at in the 1980s, the stunning facility is a veritable MOMA: a modern museum of manufacturing art.
To have any chance of one-upping China, car factories elsewhere must become hyperefficient, which includes enlisting armies of AI-controlled robots—robots that can potentially work 24/7 and never ask for a raise or a lunch break.
The factory may eventually employ 8,500 people directly, and 7,000 satellite workers, for an annual capacity of 500,000 cars—more than Tesla’s Texas Gigafactory but less than Tesla’s Shanghai plant. This past summer, just 1,340 humans were sufficient to send a constant stream of two Ioniq models down these gleaming assembly lines. The “Meta Pros” working on those lines were earning on average $58,100 a year, which is 35 percent higher than the average in Bryan County, Ga.
Clearly the days of Ford’s River Rouge complex, which employed more than 100,000 in the 1930s, are gone. As in many new factories, you’ll see surprisingly few people beyond the assembly line itself. During my visit, I spotted less than two dozen in a cavernous welding hall, where 475 robots were piecing together car chassis in a whirling, metallic dance. A steel stamping plant was so quiet that no ear protection was required, even as robots stamped out roofs and other body panels, and then stowed them in overhead racks.
Outside, human workers parked their cars beneath solar roofs that generate up to 5 percent of the plant’s electricity. Meanwhile, a fleet of 21 hydrogen fuel-cell trucks, from the Hyundai-owned Xcient, carries parts from suppliers, emitting zero tailpipe emissions. The automaker’s goal is to obtain 100 percent of the Metaplant’s energy from renewables by 2030.
An Ioniq 9 body-in-white, the basic steel skeleton of an automobile, leaves the “main buck” section of the body build line. This line is where the vehicle’s floor and sides meet to form a recognizable car. The line adapts to changing production mixes to meet customer orders, with built-in flexibility to assemble future models.Christopher Payne/Esto
Sparks fly as welding robots piece together the Ioniq 9’s “body-in-white,” the industry term for the basic steel skeleton of a car, prior to the addition of subassemblies such as the suspension, power train, body trim, and interior. The Metaplant’s welding shop houses about 500 industrial robots.Christopher Payne/Esto
Robotic welders have revolutionized car manufacturing, joining the parts of an auto body with levels of speed, precision, and safety that humans can’t match. Such advantages reduce labor costs and scrapped materials. Hyundai is also now experimenting with humanoid robots to perform welding tasks.Christopher Payne/Esto
“Body-complete” robots mount front doors onto Ioniq 5s, using machine vision and laser-measurement systems to ensure an exact fit of movable panels on each body. The robots also install mounting bolts to exact torque specifications, all validated to ensure their work meets safety and quality standards.Christopher Payne/Esto
Smart, silent robots unload trucks
When those trucks roll into docks at the Metaplant, some of the factory’s 850 robots promptly unload their parts. About 300 automated guided vehicles, or AGVs, glide silently across the factory floor with no tracks required, trained to smartly stop for humans. An AGV rolls beneath a finished Hyundai, squeezes the wheels in its robotic arms, then swiftly hoists and ferries the car where it needs to go. A companion AGV further down the line executes the exact same moves. I’ve never seen so many robotic sleds like these, or a tag team move with more efficiency and grace. Within an AI-based procurement-and-logistics system, the AGVs allocate and deliver parts to workstations for “just in time” delivery, avoiding wasted time, space, and money as they stockpile components.
An automated guided vehicle ferries dashboards for the Hyundai Ioniq 9 SUV, including each dashboard’s pair of 30-centimeter display screens. AGVs are programmed to navigate the factory, using cameras and sensors to slow or stop to avoid collisions, and emit spoken warnings to human workers in their path.Christopher Payne/Esto
“They’re delivering the right parts to the right station at the right time, so you’re no longer relying on people to make those decisions,” says Jerry Roach, senior manager of general assembly at the Metaplant.
Roach prefers that his skilled humans focus on craftsmanship, doing jobs with tactile precision that only human hands and vision can accomplish. The idea is to free people from those elements of factory work that are physically taxing, unfulfilling, and, well, robotic, so workers can use their brains and take pride in their specialized skills.
Left: Adjustable-height carriers elevate an Ioniq 5 for easy access to the central fasteners and plugs that will position suspension components and the high-voltage battery, prior to the “marriage” between the upper and lower sections of the vehicle. Those carriers provide flexibility for automated functions and manual operations by the human workers at the plant (whom Hyundai calls Meta Pros). Right: On the final assembly line, an Ioniq 9’s “top hat”—including body panels—is married to the lower “skateboard” structure, which includes the electric motors, battery, and suspension. A finished car then undergoes various tests, including a water bath to check for leaks and a quick road test outdoors. Christopher Payne/Esto
Robots, Roach says, are best tasked with heavy lifting and repetitive tasks, or those that demand digitized speed and accuracy. One example is a “collaborative” robot, sophisticated enough to work safely in close proximity to people, despite its mammoth strength. For the first time at a Hyundai factory, such a robot is installing bulky, heavy doors on the assembly line—a notoriously tricky task to perform without scratching the glossy paint or damaging surrounding panels.
Hyundai is proud of its collaborative robots, including one that can precisely install a heavy door, a tricky task for humans to perform without damaging the panels. Those robots require advanced control systems so that they can work alongside human workers without needing to be fenced off or otherwise isolated.Christopher Payne/Esto
“Guess what? Robots do that perfectly, always putting the door in the exact same place,” Roach says. “So here, that technology makes sense.”
Man’s best friend, or its mechanical counterparts, stroll the factory floor: Spot, the robotic quadrupeds from Hyundai-owned Boston Dynamics, use camera vision, sensors, and what Boston Dynamics calls “athletic intelligence” to sniff out potential welding defects.
Spot, the robot dog designed by Hyundai-owned Boston Dynamics, inspects body welds on an Ioniq 5 for defects. Equipped with a sensor suite, the quadruped bot can recharge autonomously, dynamically work around fixed or moving obstacles, and get back on its feet if it falls. Christopher Payne/Esto
Those four-legged bots may soon have a biped master: Atlas, the humanoid robot, also from Boston Dynamics. The humanoid’s physical dexterity is uncanny, with a 360-degree swiveling head that allows it to walk forward and backward without turning its body. One look at these Atlases crawling, cartwheeling, or breakdancing during testing and you might reasonably conclude they’re a potential Terminator of jobs. Hyundai executives insist that’s not the case, even as they plan to put Atlases to work in their global factories. Boston Dynamics is training these robots to sense their environments and manipulate and move parts in complex sequences.
At this backup station, high-voltage battery fasteners can be installed in an Ioniq 5. The station ensures that the assembly line keeps running even if an automated production system requires servicing. Christopher Payne/Esto
From nearby Interstate 16, Georgia drivers can see freshly painted Ioniq 5s and 9s moving along a conveyor on a windowed bridge—an intentional glimpse of what’s happening inside. They can also see their tax dollars at work, after $2.1 billion in state subsidies. Hyundai is already building a second battery plant in Georgia, and a steel plant in Louisiana, part of an expanded pledge of $21 billion in U.S. investment through 2028.
After their frames are fully welded, Ioniq 5s move along a conveyor [in the background] to an environmentally friendly paint shop. From there, the cars will travel along an elevated bridge, visible from nearby Interstate 16 in Ellabell, Ga., toward final assembly.Christopher Payne/Esto
An Ioniq 5 arrives at its final inspection station. Immediately after, a human driver gets to drive the pristine car for the first time, on a test track just outside the factory. The first Ioniq 5 rolled off the Metaplant line on 3 October 2024, with the larger Ioniq 9 kicking off production in March 2025. Christopher Payne/Esto
In a suddenly inhospitable climate for EVs, there’s nothing automatic about building and selling the cars. But Hyundai and other automakers will keep trying. They don’t have any other choice.
This article appears in the December 2025 print issue as “Inside Hyundai’s Massive Metaplant.”

Microsoft built a fake marketplace to test AI agents — they failed in surprising ways
This startup’s metal stacks could help solve AI’s massive heat problem
SoftBank, OpenAI launch new joint venture in Japan as AI deals grow ever more circular
Google Maps bakes in Gemini to improve navigation and hands-free use
Meet the Chinese Startup Using AI—and a Team of Human Workers—to Train Robots
How to Predict Biomolecular Structures Using the OpenFold3 NIM
For decades, one of biology’s deepest mysteries was how a string of amino acids folds itself into the intricate architecture of life. Researchers built...
For decades, one of biology’s deepest mysteries was how a string of amino acids folds itself into the intricate architecture of life. Researchers built painstaking simulations and statistical models, inching toward an answer but never crossing the threshold of prediction at scale. Then, deep learning changed everything. By learning the language of evolution directly from sequence data…
Google’s AI Mode gets new agentic capabilities to help book event tickets and beauty appointments
Shopify says AI traffic is up 7x since January, AI-driven orders are up 11x
Future Data Centers Could Orbit Earth, Powered by the Sun and Cooled by the Vacuum of Space
A new study suggests orbital data centers could be carbon neutral, but steep technical challenges remain.
As global demand for computing continues to explode, the carbon footprint of data centers is a growing concern. A new study outlines how hosting these facilities in space could help slash the sector’s emissions.
Data centers require enormous amounts of power and water to operate and cool the millions of chips housed within them. Current estimates from the International Energy Agency peg their electricity consumption at around 415 terawatt hours globally, roughly 1.5 percent of total consumption in 2024. And the Environmental and Energy Study Institute says that large data centers can use as much as five million gallons per day for cooling.
With demand for computing resources growing by the day, in particular since the rapid adoption of resource-guzzling generative AI across the economy, this threatens to become an unsustainable burden on the planet.
But a new paper in Nature Electronics by scientists at Nanyang Technological University in Singapore suggests that hosting data centers in space could provide a potential solution. By relying on the abundant solar energy available in orbit and releasing waste heat into the cold vacuum of space, these facilities could, in principle, become carbon neutral.

“Space offers a true sustainable environment for computing,” Wen Yonggang, lead author of the study, said in a press release. “By harnessing the sun’s energy and the cold vacuum of space, orbital data centers could transform global computing.”
To validate their proposal, the researchers used digital-twin simulations of orbital computing systems to model how they would generate power, manage heat, and maintain connectivity. The team investigated two potential architectures: one designed to reduce the footprint of data collected by satellites themselves and another that would receive data from Earth for processing.
The first model would involve integrating data processing capabilities into satellites equipped with sensors—for example, cameras for imaging the Earth. This would make it possible to carry out expensive computations on the data on board before transmitting just the results back to the ground, rather than processing the raw data in terrestrial data centers.
The other approach involves a constellation of satellites equipped with full servers that could receive data from Earth and coordinate to carry out complex computing tasks like training AI models or running large simulations. The researchers note that this kind of distributed data center architecture—as opposed to assembling a large, monolithic data center in orbit—is technologically feasible with today’s satellite and computing technologies.
The team’s analysis suggests that the considerable carbon footprint of launching hardware into space could be offset within five years of operation, after which the facilities could run indefinitely on renewable energy.
Significant technical and logistical hurdles remain. Computer chips are vulnerable to radiation, an ever-present danger in space, which would necessitate the use of specialized radiation-hardened processors. Long-term maintenance of the facilities would also require in-orbit servicing technologies that don’t yet exist. And as computing technologies rapidly improve, chips depreciate in just a few years. Keeping orbital data centers stocked with the latest and greatest could be costly.
But the NTU team isn’t the first to float the idea of shifting computing facilities into space. Last year, French defense and aerospace giant Thales published a study exploring the feasibility of the idea. And next month, the startup Starcloud will launch a satellite carrying an Nvidia H100 GPU as a first step towards creating a network of orbital data centers.
While realizing the vision is likely to require technical breakthroughs and a huge amount of investment, one solution to computing’s ever growing carbon footprint may be above our heads.
The post Future Data Centers Could Orbit Earth, Powered by the Sun and Cooled by the Vacuum of Space appeared first on SingularityHub.

Partnering with Sunflower Labs: Your Autonomous Eye in the Sky
Partnering with Sunflower Labs: Your Autonomous Eye in the Sky
With the “Beehive,” Alex, Chris, and Nick are shaping the future of real-world security.
Imagine you’re a night watchman for a self-storage facility, alone at your station at 2 a.m. A motion sensor goes off near the back fence, but a quick scan of your security cameras shows nothing. You head over to the area, calling the police on your way, but by the time you find the three cut padlocks in the last row of units, it’s too late. The thief has disappeared into the darkness—along with your customers’ belongings.
Security is a rising concern for owners of large outdoor properties, including data centers, stadiums, and warehouses, around the world. Theft, vandalism, and trespassing are costly not only in terms of repairs and insurance payments, but also reputation and lost business—and at public facilities like airports, the stakes are even higher. When you have miles of ground to patrol, cameras and human guards can only do so much, so quickly—and when you’re facing a security threat, every second counts.
Now imagine this, instead: within five seconds of that motion sensor trigger, a drone automatically takes off. It arrives at the fence long before you do—recording every second as you watch live. Humming overhead, the drone is impossible for the intruder to miss, and they flee before any damage is done.
This is the peace of mind made possible by Sunflower Labs, the company transforming how commercial campuses, critical industrial sites, and communities defend themselves. Sunflower’s “Beehive” is an AI-powered drone system that can patrol acres of properties with zero manual intervention and in almost any weather, detecting everything from people and vehicles to water leaks and fires. A new FAA approval ensures the system can operate across 99% of the U.S., keeping customers ahead of both current and planned future regulations. The system has the potential to become a core component in the future of physical security—and it’s a 10x cost advantage compared to traditional patrols.

Sunflower co-founders Alex Pachikov, Chris Eheim, and Nick de Palézieux have designed an incredible product, combining resilient hardware with easy-to-use software and integrating with a host of existing security solutions. This intentional craftsmanship is no surprise, given the technical background of the founders and their team. Many Sunflower engineers come out of ETH Zürich—one of the best universities in the world for drone and robotics technology.
Customers are as impressed as we are with both the team and product. A Swiss railway system deterred thieves, graffiti artists, and trespassers. A community in Los Angeles improved coverage for its residents. A self-storage company uses Beehive to protect their facilities, check on maintenance issues, and more. And with this round of funding, Sunflower will leverage AI to take their already exceptional product to the next level, to deepen their partnerships with other security platforms, and to bring the beehive to an even larger global audience.
The future is bright for Alex, Chris, Nick, and their growing team—and thanks to them, it is more secure for us all.
Share
Related Topics
Get the best stories from the Sequoia community.
The post Partnering with Sunflower Labs: Your Autonomous Eye in the Sky appeared first on Sequoia Capital.
Scientific frontiers of agentic AI
Scientific frontiers of agentic AI
The language AI agents might speak, sharing context without compromising privacy, modeling agentic negotiations, and understanding users commonsense policies are some of the open scientific questions that researchers in agentic AI will need to grapple with.
Conversational AI
Michael KearnsIt feels as though weve barely absorbed the rapid development and adoption of generative AI technologies such as large language models (LLMs) before the next phenomenon is already upon us, namely agentic AI. Standalone LLMs can be thought of as chatbots in a sandbox, the sandbox being a metaphor for a safe and contained play space with limited interaction with the world beyond. In contrast, the vision of agentic AI is a near (or already here?) future in which LLMs are the underlying engines for complex systems that have access to rich external resources such as consumer apps and services, social media, banking and payment systems in principle, anything you can reach on the Internet. A dream of the AI industry for decades, the agent of agentic AI is an intelligent personal assistant that knows your goals and preferences and that you trust to act on your behalf in the real world, much as you might a human assistant.
For example, in service of arranging travel plans, my personal agentic AI assistant would know my preferences (both professional and recreational) for flights and airlines, lodging, car rentals, dining, and activities. It would know my calendar and thus be able to schedule around other commitments. It would know my frequent-flier numbers and hospitality accounts and be able to book and pay for itineraries on my behalf. Most importantly, it would not simply automate these tasks but do so intelligently and intuitively, making obvious decisions unilaterally and quietly but being sure to check in with me whenever ambiguity or nuance arises (such as whether those theater tickets on a business trip to New York should be charged to my personal or work credit card).
To AI insiders, the progression from generative to agentic AI is exciting but also natural. In just a few years, we have gone from impressive but glorified chatbots with myriad identifiable shortcomings to feature-rich systems exhibiting human-like capabilities not only in language and image generation but in coding, mathematical reasoning, optimization, workflow planning, and many other areas. The increased skill set and reliability of core LLMs has naturally caused the industry to move up the stack, to a world in which the LLM itself fades into the background and becomes a new kind of intelligent operating system upon which all manner of powerful functionality can be built. In the same way that your PC or Mac seamlessly handles many details that the vast majority of users dont (want to) know about like exactly how and where on your hard drive to store and find files, the networking details of connecting to remote web servers, and other fine-grained operating-system details agentic systems strive to abstract away the messy and tedious details of many higher-level tasks that, today, we all perform ourselves.
But while the overarching vision of agentic AI is already relatively clear, there are some fundamental scientific and technical questions about the technology whose answers or even proper formulation are uncertain (but interesting!). Well explore some of them here.
What language will agents speak?
The history of computing technology features a steady march toward systems and devices that are ever more friendly, accessible, and intuitive to human users. Examples include the gradual displacement of clunky teletype monitors and obscure command-line incantations by graphical user interfaces with desktop and folder metaphors, and the evolution from low-level networked file transfer protocols to the seamless ease of the web. And generative AI itself has also made previously specialized tasks like coding accessible to a much broader base of users. In other words, modern technology is human-centric, designed for use and consumption by ordinary people with little or no specialized training.
But now these same technologies and systems will also need to be navigated by agentic AI, and as adept as LLMs are with human language, it may not be their most natural mode of communication and understanding. Thus, a parallel migration to the native language of generative AI may be coming.
What is that native language? When generative AI consumes a piece of content whether it be a user prompt, a document, or an image it translates it into an internal representation that is more convenient for subsequent processing and manipulation. There are many examples in biology of such internal representations. For instance, in our own visual systems, it has been known for some time that certain types of inputs (such as facial images) cause specific cells in our brains to respond (a phenomenon known as neuronal selectivity). Thus, an entire category of important images elicits similar neural behaviors.
In a similar vein, the neural networks underlying modern AI typically translate any input into what is known as an embedding space, which can be thought of as a physical map in which items with similar meanings are placed near each other, and those with unrelated meanings are placed far apart. For example, in an image-embedding space, two photos of different families would be nearer to each other than either would be to a landscape. In a language-embedding space, two romance novels would be nearer to each other than to a car owners manual. And hybrid or multimodal embedding spaces would place images of cars near their owner manuals.
Embeddings are an abstraction that provides great power and generality, in the form of the ability to represent not the literal original content (like a long sequence of words) but something closer to its underlying meaning. The price for this abstraction is loss of detail and information. For instance, the embedding of this entire article would place it in close proximity to similar content (for instance, general-audience science prose) but would not contain enough information to re-create the article verbatim. The lossy nature of embeddings has implications we shall return to shortly.
Embeddings are learned from the massive amount of information on the Internet and elsewhere about implicit correspondences. Even aliens landing on earth who could read English but knew nothing else about the world would quickly realize that doctor and hospital are closely related because of their frequent proximity in text, even if they had no idea what these words actually signified. Furthermore, not only do embeddings permit generative AI to understand existing content, but they allow it to generate new content. When we ask for a picture of a squirrel on a snowboard in the style of Andy Warhol, it is the embedding that lets the technology explore novel images that interpolate between those of actual Warhols, squirrels, and snowboards.
Thus, the inherent language of generative (and therefore agentic) AI is not the sentences and images we are so familiar with but their embeddings. Let us now reconsider a world in which agents interact with humans, content, and other agents. Obviously, we will continue to expect agentic AI to communicate with humans in ordinary language and images. But there is no reason for agent-to-agent communication to take place in human languages; per the discussion above, it would be more natural for it to occur in the native embedding language of the underlying neural networks.
My personal agent, working on a vacation itinerary, might ingest materials such as my previous flights, hotels, and vacation photos to understand my interests and preferences. But to communicate those preferences to another agent say, an agent aggregating hotel details, prices, and availability it will not provide the raw source materials; in addition to being massively inefficient and redundant, that could present privacy concerns (more on this below). Rather, my agent will summarize my preferences as a point, or perhaps many points, in an embedding space.
By similar reasoning, we might also expect the gradual development of an agentic Web meant for navigation by AI, in which the text and images on websites are pre-translated into embeddings that are illegible to humans but are massively more efficient than requiring agents to perform these translations themselves with every visit. In the same way that many websites today have options for English, Spanish, Chinese, and many other languages, there would be an option for Agentic.
All the above presupposes that embedding spaces are shared and standardized across generative and agentic AI systems. This is not true today: embeddings differ from model to model and are often considered proprietary. Its as if all generative AI systems speak slightly different dialects of some underlying lingua franca. But these observations about agentic language and communication may foreshadow the need for AI scientists to work toward standardization, at least in some form. Each agent can have some special and proprietary details to its embeddings for instance, a financial-services agent might want to use more of its embedding space for financial terminology than an agentic travel assistant would but the benefits of a common base embedding are compelling.
Keeping things in context
Even casual users of LLMs may be aware of the notion of context, which is informally what and how much the LLM remembers and understands about its recent interactions and is typically measured (at least cosmetically) by the number of words or tokens (word parts) recalled. There is again an apt metaphor with human cognition, in the sense that context can be thought of as the working memory of the LLM. And like our own working memory, it can be selective and imperfect.
If we participate in an experiment to test how many random digits or words we can memorize at different time scales, we will of course eventually make mistakes if asked to remember too many things for too long. But we will not forget what the task itself is; our short-term memory may be fallible, but we generally grasp the bigger picture.
These same properties broadly hold for LLM context which is sometimes surprising to users, since we expect computers to be perfect at memorization but highly fallible on more abstract tasks. But when we remember that LLMs do not operate directly on the sequence of words or tokens in the context but on the lossy embedding of that sequence, these properties become less mysterious (though perhaps not less frustrating when an LLM cant remember something it did just a few steps ago).
Some of the principal advances in LLM technology have been around improvements in context: LLMs can now remember and understand more context and leverage that context to tailor their responses with greater accuracy and sophistication. This greater window of working memory is crucial for many tasks to which we would like to apply agentic AI, such as having an LLM read and understand the entire code base of a large software development project, or all the documents relevant to a complex legal case, and then be able to reason about the contents.
How will context and its limitations affect agentic AI? If embeddings are the language of LLMs, and context is the expression of an LLMs working memory in that language, a crucial design decision in agent-agent interactions will be how much context to share. Sharing too little will handicap the functionality and efficiency of agentic dialogues; sharing too much will result in unnecessary complexity and potential privacy concerns (just as in human-to-human interactions).
Let us illustrate by returning to my personal agent, who having found and booked my hotel is working with an external airline flight aggregation agent. It would be natural for my agent to communicate lots of context about my travel preferences, perhaps including conditions under which I might be willing to pay or use miles for an upgrade to business class (such as an overnight international flight). But my agent should not communicate context about my broader financial status (savings, debt, investment portfolio), even though in theory these details might correlate with my willingness to pay for an upgrade. When we consider that context is not my verbatim history with my travel agent, but an abstract summary in embedding space, decisions about contextual boundaries and how to enforce them become difficult.
Indeed, this is a relatively untouched scientific topic, and researchers are only just beginning to consider questions such as what can be reverse-engineered about raw data given only its embedding. While human or system prompts to shape inter-agent dealings might be a stopgap (be sure not to tell the flight agent any unnecessary financial information), a principled understanding of embedding privacy vulnerabilities and how to mitigate them (perhaps via techniques such as differential privacy) is likely to be an important research area going forward.
Agentic bargains
So far, weve talked a fair amount about interagent dialogues but have treated these conversations rather generally, much as if we were speaking about two humans in a collaborative setting. But there will be important categories of interaction that will need to be more structured and formal, with identifiable outcomes that all parties commit to. Negotiation, bargaining, and other strategic interactions are a prime example.
I obviously want my personal agent, when booking hotels and flights for my trips, to get the best possible prices and other conditions (room type and view, flight seat location, and so on). The agents aggregating hotels and flights would similarly prefer that I pay more rather than less, on behalf of their own clients and users.
For my agent to act in my interests in these settings, Ill need to specify at least some broad constraints on my preferences and willingness to pay for them, and not in fuzzy terms: I cant expect my agent to simply know a bargain when it sees one the way I might if I were handling all the arrangements myself, especially because my notion of a bargain might be highly subjective and dependent on many factors. Again, a near-term makeshift approach might address this via prompt shaping be sure to get the best deal possible, as long as the flight is nonstop and leaves in the morning, and I have an aisle seat but longer-term solutions will have to be more sophisticated and granular.
Of course, the mathematical and scientific foundations of negotiating and bargaining have been well studied for decades by game theorists, microeconomists, and related research communities. Their analyses typically begin by presuming the articulation of utility functions for all the parties involved an abstraction capturing (for example) my travel preferences and willingness to pay for them. The literature also considers settings in which I cant quantitatively express my own utilities but know bargains when I see them, in the sense that given two options (a middle seat on a long flight for $200 vs. a first-class seat for $2,000), I will make the choice consistent with my unknown utilities. (This is the domain of the aptly named utility elicitation.)
Much of the science in such areas is devoted to the question of what should happen when fully rational parties with precisely specified utilities, perfect memory, and unlimited computational power come to the proverbial bargaining table; equilibrium analysis in game theory is just one example of this kind of research. But given our observations about the human-like cognitive abilities and shortcomings of LLMs, perhaps a more relevant starting point for agentic negotiation is the field of behavioral economics. Instead of asking what should happen when perfectly rational agents interact, behavioral economics asks what does happen when actual human agents interact strategically. And this is often quite different, in interesting ways, than what fully rational agents would do.
For instance, consider the canonical example of behavioral game theory known as the ultimatum game. In this game, there is $10 to potentially divide between two players, Alice and Bob. Alice first proposes any split she likes. Bob then either accepts Alices proposal, in which case both parties get their proposed shares, or rejects Alices proposal, in which case each party receives nothing. The equilibrium analysis is straightforward: Alice, being fully rational and knowing that Bob is also, proposes the smallest nonzero amount to Bob, which is a penny. Bob, being fully rational, would prefer to receive a penny than nothing, so he accepts.
Nothing remotely like this happens when humans play. Across hundreds of experiments varying myriad conditions social, cultural, gender, wealth, etc. a remarkably consistent aggregate behavior emerges. Alice almost always proposes a share to Bob of between $3 and $5 (the fact that Alice gets to move first seems to prime both players for Bob to potentially get less than half the pie). And conditioned on Alices proposal being in this range, Bob almost always accepts her offer. But on those rare occasions in which Alice is more aggressive and offers Bob an amount much less than $3, Bobs rejection rate skyrockets. Its as if pairs of people who have never heard of or played the ultimatum game before have an evolutionarily hardwired sense of whats fair in this setting.
Now back to LLMs and agentic AI. There is already a small but growing literature on what we might call LLM behavioral game theory and economics, in which experiments like the one above are replicated except human participants are replaced by AI. One early work showed that LLMs almost exactly replicated human behavior in the ultimatum game, as well as other classical behavioral-economics findings.
Note that it is possible to simulate the demographic variability of human subjects in such experiments via LLM prompting, e.g., You are Alice, a 37-year-old Hispanic medical technician living in Boston, Massachusetts. Other studies have again shown human-like behavior of LLMs in trading games, price negotiations, and other settings. A very recent study claims that LLMs can even engage in collusive price-fixing behaviors and discusses potential regulatory implications for AI agents.
Once we have a grasp on the behaviors of agentic AI in strategic settings, we can turn to shaping that behavior in desired ways. The field of mechanism design in economics complements areas like game theory by asking questions like given that this is how agents generally negotiate, how can we structure those negotiations to make them fair and beneficial? A classic example is the so-called second-price auction, where the highest bidder wins the item but only pays the second highest bid. This design is more truthful than a standard first-price auction, in the sense that everyones optimal strategy is to simply bid the price at which they are indifferent to winning or losing (their subjective valuation of the item); nobody needs to think about other agents behaviors or valuations.
We anticipate a proliferation of research on topics like these, as agentic bargaining becomes commonplace and an important component of what we delegate to our AI assistants.
The enduring challenge of common sense
Ill close with some thoughts on a topic that has bedeviled AI from its earliest days and will continue to do so in the agentic era, albeit in new and more personalized ways. Its a topic that is as fundamental as it is hard to define: common sense.
By common sense, we mean things that are obvious, that any human with enough experience in the world would know without explicitly being told. For example, imagine a glass full of water sitting on a table. We would all agree that if we move the glass to the left or right on the table, its still a glass of water. But if we turn it upside down, its still a glass on the table, but no longer a glass of water (and is also a mess to be cleaned up). Its quite unlikely any of us were ever sat down and run through this narrative, and its also a good bet that youve never deliberately considered such facts before. But we all know and agree on them.
Figuring out how to imbue AI models and systems with common sense has been a priority of AI research for decades. Before the advent of modern large-scale machine learning, there were efforts like the Cyc project (for encyclopedia), part of which was devoted to manually constructing a database of commonsense facts like the ones above about glasses, tables, and water. Eventually the consumer Internet generated enough language and visual data that many such general commonsense facts could be learned or inferred: show a neural network millions of pictures of glasses, tables and water and it will figure things out. Very early research also demonstrated that it was possible to directly encode certain invariances (similar to shifting a glass of water on a table) into the network architecture, and LLM architectures are similarly carefully designed in the modern era.
But in agentic AI, we expect our proxies to understand not only generic commonsense facts of the type weve been discussing but also common sense particular to our own preferences things that would make sense to most people if only they understood our contexts and perspectives. Here a pure machine learning approach will likely not suffice. There just wont be enough data to learn from scratch my subjective version of common sense.
For example, consider your own behavior or policy around leaving doors open or closed, locked or unlocked. If youre like me, these policies can be surprisingly nuanced, even though I follow them without thought all the time. Often, I will close and lock doors behind me for instance, when I leave my car or my house (unless Im just stepping right outside to water the plants). Other times I will leave a door unlocked and open, such as when Im in my office and want to signal I am available to chat with colleagues or students. I might close but leave unlocked that same door when I need to focus on something or take a call. And sometimes Ill leave my office door unlocked and open even when Im not in it, despite there being valuables present, because I trust the people on my floor and Im going to be nearby.
We might call behaviors like these subjective common sense, because to me they are natural and obvious and have good reasons behind them, even though I follow them almost instinctually, the same way I know not to turn a glass of water upside down on the table. But you of course might have very different behaviors or policies in the same or similar situations, with your own good reasons.
The point is that even an apparently simple matter like my behavior regarding doors and locks can be difficult to articulate. But agentic AI will need specifications like this: simply replace doors with online accounts and services and locks with passwords and other authentication credentials. Sometimes we might share passwords with family or friends for less-critical privacy-sensitive resources like Netflix or Spotify, but we would not do the same for bank accounts and medical records. I might be less rigorous about restricting access to, or even encrypting, the files on my laptop than I would be about files I store in the cloud.
The circumstances under which I trust my own or other agents with resources that need to be private and secure will be at least as complex as those regarding door closing and locking. The primary difficulty is not in having the right language or formalisms to specify such policies: there are good proposals for such specification frameworks and even for proving the correctness of their behaviors. The problem is in helping people articulate and translate their subjective common sense into these frameworks in the first place.
Conclusion
The agentic-AI era is in its infancy, but we should not take that to mean we have a long and slow development and adoption period before us. We need only look at the trajectory of the underlying generative AI technology from being almost entirely unknown outside of research circles as recently as early 2022 to now being arguably the single most important scientific innovation of the century so far. And indeed, there is already widespread use of what we might consider early agentic systems, such as the latest coding agents.
Far beyond the initial autocomplete for Python tools of a few years ago, such agents now do so much more writing working code from natural-language prompts and descriptions, accessing external resources and datasets, proactively designing experiments and visualizing the results, and most importantly (especially for a novice programmer like me), seamlessly handling the endless complexity of environment settings, software package installs and dependencies, and the like. My Amazon Scholar and University of Pennsylvania colleague Aaron Roth and I recently wrote a machine learning paper of almost 50 pages complete with detailed definitions, theorem statements and proofs, code, and experiments using nothing except (sometimes detailed) English prompts to such a tool, along with expository text we wrote directly. This would have been unthinkable just a year ago.
Despite the speed with which generative AI has permeated industry and society at large, its scientific underpinnings go back many decades, arguably to the birth of AI but certainly no later than the development of neural-network theory and practice in the 1980s. Agentic AI built on top of these generative foundations, but quite distinct in its ambitions and challenges has no such deep scientific substrate on which to systematically build. Its all quite fresh territory. Ive tried to anticipate some of the more fundamental challenges here, and Ive probably got half of them wrong. To paraphrase the Philadelphia department store magnate John Wanamaker, I just dont know which half yet.
Research areas: Conversational AI, Machine learning
Tags: Agentic AI, Generative AI, Large language models (LLMs)
OpenAI Signs $38 Billion Deal With Amazon
AI’s Next Frontier? An Algorithm for Consciousness