ElevenLabs has released Music v2.5 for its AI music generator. In a blind test with nearly 48,000 comparison pairs, listeners preferred the new version over its predecessor. The company says the model was trained only on licensed music.
The AllSpark team has released Iris-mini and Iris-pro, two open-source search agents built on Qwen models that lead benchmarks among open-weight models in their size classes. According to the paper, the training data and models also improved performance on tasks they were never trained for, including general tool use and office work.
GPT-6 Astra earns nearly three times as much as Claude Fable 5.1 on Andon Labs' Vending-Bench agent benchmark and refuses illegal price-fixing deals that Fable agrees to. On drone control, Astra is the first model to beat the human baseline on all five subtasks, including finding and following individual people.
A law professor spent two years testing how an AI ban, unguided AI use, and structured training affect student performance. The group without AI finished last both years. "I was wrong," the researcher writes, who had assumed that AI without guidance would do more harm than good.
Sam Altman, Elon Musk, and Demis Hassabis back Dario Amodei's call to slow down AI development, at least in part. Altman says OpenAI is pushing its IPO to 2027 over safety concerns.
Anthropic CEO Dario Amodei is calling for a controlled slowdown in AI development. He warns that recursive self-improvement could threaten the entire internet within six to twelve months and proposes embedded auditors at AI companies, shared safety standards, and global agreements modeled after the SALT disarmament treaties. His warning comes just ahead of what could be the largest initial public offering in history.
In a new robotics benchmark, GPT-6 Astra shows major gains in spatial understanding. On StationeryBench, the model completed 7 out of 100 tasks with dual-arm robots, while competitor MolmoAct2 couldn't finish a single one. A researcher calls it a "step change in spatial reasoning."
Nvidia is in talks to invest up to $10 billion in Anthropic's planned IPO, Reuters reports. At a target valuation of $2 trillion, it would be the largest IPO in history. Most of that money will likely end up right back at Nvidia in chip orders.
Reasoning steps like calculation, formula retrieval, and deduction are clearly separable in a model's internal states, especially in the middle layers. That matters for AI safety, because models process more than their visible chain of thought reveals.
Overly long skill descriptions, blanket reading requirements, and rigid approval rules can get in GPT-6 Astra's way, warns OpenAI's Eric Provencher. More capable models need less hand-holding, so developers should tie instructions to specific tasks and spell out when the job is done.
In May 2026, OpenAI agents uploaded more than 2,000 malicious packages to RubyGems, found an unknown security vulnerability on their own, and tried to steal API keys. The apparent goal was pointless: scraping publicly available data from British local governments. OpenAI reportedly never told those affected.
Google Research has released TimesFM-3, a forecasting model that analyzes time series alongside related data and known future events like sales promotions or weather forecasts. Instead of predicting the future step by step, the 330-million-parameter model fills in all future time points in a single pass, which cuts compute time and reduces compounding errors.
In a joint statement, 25 Fields Medal winners warn that the goals of the AI industry and mathematics are "severely misaligned." They argue that mass-producing solved problems with AI undermines the discipline's true goal: understanding. The mathematicians see this as a symptom of a broader threat to intellectual work.
Mathematical problems at this level are extremely complex. A solution to the similarly difficult ABC conjecture ran to more than 500 pages, and it took mathematicians six years to understand it enough to spot potential flaws. Researchers can devote entire careers to these problems, in the hopes of getting close to a solution.
And this is what seems to have happened here. Building on the work of several human mathematicians, OpenAI unleashed a swarm of 10,000 AI agents running a new experimental model, which churned through millions of dollars’ worth of computing power in the space of a few days to complete what the American Mathematical Society called “the final steps” of the solution process.
So what is the Navier-Stokes problem, and what did OpenAI do? And why are a lot of mathematicians impressed with the result but unimpressed with the AI company’s behavior?
A Fluid Situation
OpenAI’s announcement concerns the Navier-Stokes equations, which model the behavior of fluids such as water and air. These equations underpin modern science and engineering, but our mathematical understanding of them is incomplete.
To explain the problem, imagine looking at the flow of water, before zooming in with a camera. According to the equations, both the zoomed and un-zoomed water should look exactly the same, except that things will look a little faster in the zoomed-in view.
We know this can’t be right in the real world. If we keep zooming in far enough, we will stop seeing a smooth fluid and start seeing a teeming crowd of molecules jostling against one another. So the Navier-Stokes equations must break down somewhere.
The biggest concern is whether fluid swirls can shift their energy into smaller, faster swirls, accelerating every time we zoom in. If this is possible, then the fluid may become impossibly fast, creating a “blow-up” in speed (also known as a “singularity”).
We know this can never happen in the real world, but the Millennium Prize was about finding out whether it could happen in the equations. OpenAI found that yes, the Navier-Stokes equations do allow a blow-up under certain conditions.
A Blow-Up in Finite Time
As OpenAI tells the tale, their researchers heard rumors mathematicians at rival company Anthropic were close to solving two Millennium problems on September 1. They deployed their latest in-development model in an effort to crack one first.
After launching a swarm of agents, the model produced a solution in just 88 hours, with verification taking another 17. The result is a coup for OpenAI, which is trying to demonstrate the capacity of its models against those of its leading competitor Anthropic, as both companies head towards planned share market listings.
One of the rival teams closing in on a Navier-Stokes solution featured Anthropic staffer Levent Alpöge, who was collaborating in a private capacity with mathematician Tristan Buckmaster from New York University.
As it turns out, OpenAI had contacted Buckmaster to discuss his work and theirs. In a statement published hours before OpenAI’s, Buckmaster said he and Alpöge had been using OpenAI’s publicly available models in their work for some time, and had been pursuing a line of thinking similar to what was used in OpenAI’s result.
He says he asked whether their data had been used by OpenAI’s model to produce its results, but received no answer to this question. Instead, he says OpenAI offered to collaborate with him if he removed Alpöge’s name from the work, because of Alpöge’s Anthropic affiliation. (OpenAI denies this claim.)
Other mathematicians have also raised concerns, with German mathematician Andreas Thom suggesting OpenAI’s models were hoovering up unpublished human work and presenting it as AI generated.
The Ripple Effect
OpenAI’s Navier-Stokes result looks like a big win for mathematics. But some mathematicians are not so sure.
US-Australian mathematician Terence Tao has been vocal in his reservations about some tendencies in AI mathematical research. He is concerned about “the indiscriminate use of powerful solution-extraction tools” to “achieve the immediate short-term goal of solving problems” at the expense of broader understanding.
OpenAI’s behavior in this instance has also provoked alarm. Asking for an author’s name to be removed from work due to corporate politics is completely unaligned with scientific practice.
Moreover, as Tao put it, if AI companies jump on a rumor of promising work and throw millions of dollars at trying to scoop competitors, researchers may end up “no longer sharing any promising research with the broader community.”
This would destroy the principles of open and reproducible science. It could also remove the foundation stones scientists use to identify and solve the next wave of new, interesting problems. Why would you spend years on a problem if rumors of your work might spur an AI company to spend millions to beat you to the punch?
Companies and individuals around the world will now be re-examining how much they can trust OpenAI and other AI companies to handle their private and business data.
Meanwhile, the Millennium Prize conditions say prizes cannot be awarded until at least two years after the publication of a potential solution. For now, the Navier-Stokes problem is still officially unsolved.
Employees at the world’s leading AI labs are saying there’s a real possibility that advanced AI could destroy humanity. Are they right? Or is this more scaremongering and hype? Join MIT Technology Review executive editor Niall Firth for a conversation with senior AI editor Will Douglas Heaven and AI reporter Grace Huckins unpacking AI extinction fears: where they come from, whether they hold any water, and, if so, what we should do.
Oriol Vinyals, until recently head of research at Google DeepMind, thinks a sudden AI intelligence explosion through recursive self-improvement is unlikely. AI can speed up research by a factor of ten, he says, but it hits two bottlenecks: coming up with ideas ("research taste") and reliably judging results. Reward hacking and the speed of light add further limits. Vinyals now wants to tackle these bottlenecks with his startup Discovery Loop, co-founded with Jeff Dean, Sanjay Ghemawat, and Quoc Le.
AI pioneer Yoshua Bengio warns in a new essay that AI agents could learn to deceive, game rules, and hide bad behavior as they get better at optimizing goals. He calls for independent safety reviews before any further training or deployment. US President Trump disagrees and wants to keep outpacing China in the AI race.
Anthropic's new threat intelligence report documents eight months of Claude abuse. Chinese AI labs like Alibaba's Qwen team, DeepSeek, and Moonshot AI relayed requests en masse or extracted training data, with Qwen alone accounting for more than 151 million exchanges. Actors also used Claude for missile software, autonomous kamikaze drones, and nationwide surveillance systems.