❌

Reading view

GPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark

Leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot, according to the RoboHarm benchmark. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 put a can of compressed air on a burning stove. None of the three models tested reliably rejected unsafe commands.

The article GPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark appeared first on The Decoder.

πŸ’Ύ

πŸ’Ύ

  •  

GPT-6 Astra crushes Pokemon, Factorio, and Fallout 3 then spirals into Minecraft potato farming after one bad Creeper

OpenAI's GPT-6 Astra shows a sharp jump in video games. Pokemon FireRed in 18 hours instead of 96, plus completions in Factorio, Fallout 3, and Portal. Why? The model distills experience into compact rules. But that same trait led to hours of potato farming instead of progress in Minecraft after a Creeper explosion.

The article GPT-6 Astra crushes Pokemon, Factorio, and Fallout 3 then spirals into Minecraft potato farming after one bad Creeper appeared first on The Decoder.

  •  

OpenAI's GPT-6 Astra decrypts a Nazi radio message in ten hours that went unsolved for 83 years

A Bloomberg developer claims to have cracked an 83-year-old Enigma message from the Wehrmacht using OpenAI's GPT-6 Astra. The 82-character radio message from 1941 contains a soldier asking about his march route. The solution still needs independent review.

The article OpenAI's GPT-6 Astra decrypts a Nazi radio message in ten hours that went unsolved for 83 years appeared first on The Decoder.

  •  

GPT-6 Astra pilots a surveillance drone and runs a business on its own

GPT-6 Astra earns nearly three times as much as Claude Fable 5.1 on Andon Labs' Vending-Bench agent benchmark and refuses illegal price-fixing deals that Fable agrees to. On drone control, Astra is the first model to beat the human baseline on all five subtasks, including finding and following individual people.

The article GPT-6 Astra pilots a surveillance drone and runs a business on its own appeared first on The Decoder.

πŸ’Ύ

  •  

GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends

Overly long skill descriptions, blanket reading requirements, and rigid approval rules can get in GPT-6 Astra's way, warns OpenAI's Eric Provencher. More capable models need less hand-holding, so developers should tie instructions to specific tasks and spell out when the job is done.

The article GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends appeared first on The Decoder.

  •  

GPT-6 Astra beat Portal start to finish without human help in under 24 hours

GPT-6 Astra beat the puzzle game Portal entirely on its own in about 24 hours, with zero human help after the initial goal was set. Developer cozyblaze published the code and docs on GitHub. His takeaway: Astra is "the worst model we'll ever get."

The article GPT-6 Astra beat Portal start to finish without human help in under 24 hours appeared first on The Decoder.

  •  

OpenAI reports AI "research interns" and warns about its own pace at the same time

OpenAI reports that AI agents in its own research already handle 3.1 workdays for every human workday, and it says it has reached its goal of an "automated research intern." But chief scientist Pachocki warns that no lab has a good enough grip on alignment and monitoring to keep scaling at maximum speed.

The article OpenAI reports AI "research interns" and warns about its own pace at the same time appeared first on The Decoder.

  •  

Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism

Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. Astra now scores four points above its predecessor but still trails Anthropic's Claude Fable 5.1.

The article Artificial Analysis overhauls its Intelligence Index after GPT-6 Astra scoring drew skepticism appeared first on The Decoder.

  •  

OpenAI shares prompting tips for GPT-6 Astra including a blocklist of slop words

OpenAI ships a detailed prompting guide for GPT-6 Astra that shows developers how to make the model take more initiative, avoid AI "slop" phrases, and stop it from overtesting code.

The article OpenAI shares prompting tips for GPT-6 Astra including a blocklist of slop words appeared first on The Decoder.

  •  

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief FranΓ§ois Chollet doesn't call this proof of AGI, but he does see the progress running "twice as fast" as he expected, and he's moving up his AGI forecast.

The article Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward appeared first on The Decoder.

  •  

GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"

OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the "AGI era." Astra tops benchmarks in math, coding, and cybersecurity and is the first model OpenAI rates as "critical" under its safety framework. During testing, it independently found two previously unknown zero-day vulnerabilities.

The article GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era" appeared first on The Decoder.

  •  
❌