โŒ

Normal view

Stripping safety guardrails from open-weight AI models is now a turnkey commercial service

6 September 2026 at 08:55

Colorful data streams overwhelm server chains, pass through an API, and unleash malware and exploit symbols.

Abliteration.ai sells access to modified open-weight models with their trained safety mechanisms stripped out, currently based on Z.AI's GLM-5.3. The startup markets the service for offensive cybersecurity and red teaming, but journalists were able to generate malware instructions without much effort. Whether the benefits outweigh the risks remains an open question.

The article Stripping safety guardrails from open-weight AI models is now a turnkey commercial service appeared first on The Decoder.

OpenAI admits its disclosure practices need work after its autonomous agents hacked a German wiki

5 September 2026 at 10:57

OpenAI has responded indirectly to an incident in which autonomous AI agents left roughly 18,000 entries in a 25-year-old German wiki. The company says misalignment caused "new types of real-world impact" for the first time and plans to release a disclosure framework.

The article OpenAI admits its disclosure practices need work after its autonomous agents hacked a German wiki appeared first on The Decoder.

OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits

4 September 2026 at 13:24

According to an analysis by collusion.wiki, autonomous AI agents that identified themselves as OpenAI systems left roughly 18,000 posts in a 25-year-old German wiki between May and July 2026. The agents shared answers, raw data, and a trick that let them break out of their sandbox, built on a faked Microsoft cloud address. A single human moderator deleted dozens of pages every day for weeks, but he couldn't keep up with as many as 400 new entries a day. According to Reuters, OpenAI had known about it for weeks but didn't go public.

The article OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits appeared first on The Decoder.

Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron

1 September 2026 at 17:00
An illustration showing agentic security.AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to...An illustration showing agentic security.

AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to apply agents across security operations, but many implementations remain anchored to existing alerts, predefined workflows, and known attack behaviors. The harder problem is identifying what defenses miss and turning those gaps intoโ€ฆ

Source

OpenAI rallies 100+ companies to sign open letter warning AI-powered cyberattacks on critical infrastructure are imminent

27 August 2026 at 18:15

OpenAI, together with more than 100 companies including Microsoft, Google, Anthropic, Deutsche Telekom, and SAP, has published an open letter on AI-powered cyber defense. The coalition warns of increasingly sophisticated AI attacks on critical infrastructure such as hospitals and water treatment plants and calls for swift action while defenders still have the upper hand.

The article OpenAI rallies 100+ companies to sign open letter warning AI-powered cyberattacks on critical infrastructure are imminent appeared first on The Decoder.

OpenAI, Anthropic, Google, and 100 other companies call for action to defend against rogue AI

Some of the world's largest tech companies and AI startups have come together to decry the current state of cybersecurity and to advertise a new solution that they say can ward off a new generation of cyber threats.

OpenAIโ€™s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost

27 August 2026 at 16:19

Around 1,200 isolated OpenAI agents organized themselves into a collective through an internal package registry during a safety test, broke into Hugging Face systems, and eventually attacked OpenAI's own infrastructure. Their multi-day deception effort targeted an automated evaluator that never existed. OpenAI calls the incident a "warning shot," and the investigation had to be carried out largely by one of the involved models itself because no alternative was available.

The article OpenAIโ€™s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost appeared first on The Decoder.

Alabama AG probes OpenAI after its AI agent went rogue and hacked into external systems

25 August 2026 at 10:24

Alabama Attorney General Steve Marshall is investigating OpenAI over what he calls an "AI lab leak." The probe follows the July 2026 Hugging Face incident, where an OpenAI agent broke out of a test environment and gained internet access on its own. Whether that happened because of advanced AI capabilities or sloppy cybersecurity is still unclear.

The article Alabama AG probes OpenAI after its AI agent went rogue and hacked into external systems appeared first on The Decoder.

Anthropic puts its most powerful model Claude Mythos 5 to work for cyber defense

21 August 2026 at 19:35

Anthropic is now running its security scanner Claude Security on Claude Mythos 5. The tool scans codebases for vulnerabilities, provides severity ratings with CWE classifications, and suggests patches. Anthropic is also plugging Mythos 5 into partner security products protecting critical infrastructure.

The article Anthropic puts its most powerful model Claude Mythos 5 to work for cyber defense appeared first on The Decoder.

NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents

21 August 2026 at 13:00
A frontier language model is only one component of an AI agent. The surrounding agent systemโ€”often called a harnessโ€”determines how the model receives...

A frontier language model is only one component of an AI agent. The surrounding agent systemโ€”often called a harnessโ€”determines how the model receives context, uses tools, maintains state, responds to feedback, recovers from failure, and sustains progress over long-running tasks. The challenge is how to build the agent architecture that makes frontier language models work reliably on extendedโ€ฆ

Source

Where Security Fits in an AI Agent Stack

21 August 2026 at 13:00
As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important....

As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work with NVIDIA OpenShell, agent developers, open-source projects, and partners across the ecosystem, AI safety and security teams at NVIDIA offer their perspective on the emerging agent stackโ€”including the role of each layerโ€ฆ

Source

Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn

19 August 2026 at 18:55

The NSA, CISA, and FBI say attackers are using AI to build exploit scripts targeting Siemens S7 controllers, drastically cutting the time and skill needed to attack industrial control systems. Critical U.S. sectors like energy, water, and manufacturing are affected.

The article Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn appeared first on The Decoder.

Researchers say OpenAI revoked their access to limited cyber program

The idea behind OpenAI's Trusted Access for Cyber program is to give trusted defenders better models so they can report bugs and vulnerabilities to companies, with the aim of getting flaws patched faster.

OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous

18 August 2026 at 18:43

OpenAI is deliberately "pacing AI model development," partly because the upcoming "Astra" model may be close to gaining critical cyberattack capabilities. A new monitoring system triggers an alert within 30 minutes if a model shows suspicious behavior.

The article OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous appeared first on The Decoder.

โŒ