❌

Normal view

Frontier AI labs still won’t say how they’d contain a rogue model

A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as AI systems increasingly demonstrate unexpected and potentially dangerous behavior.

Psychological methods reveal major weaknesses in AI security testing

22 August 2026 at 07:00

A hardened AI hardware module in a glass enclosure is tested for security vulnerabilities using electrical and physical probes.

Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to day. The study also offers a method for catching models that act more cautious during tests than they do in normal use.

The article Psychological methods reveal major weaknesses in AI security testing appeared first on The Decoder.

❌