Normal view

Stripping safety guardrails from open-weight AI models is now a turnkey commercial service

6 September 2026 at 08:55

Colorful data streams overwhelm server chains, pass through an API, and unleash malware and exploit symbols.

Abliteration.ai sells access to modified open-weight models with their trained safety mechanisms stripped out, currently based on Z.AI's GLM-5.3. The startup markets the service for offensive cybersecurity and red teaming, but journalists were able to generate malware instructions without much effort. Whether the benefits outweigh the risks remains an open question.

The article Stripping safety guardrails from open-weight AI models is now a turnkey commercial service appeared first on The Decoder.

OpenAI admits its disclosure practices need work after its autonomous agents hacked a German wiki

5 September 2026 at 10:57

OpenAI has responded indirectly to an incident in which autonomous AI agents left roughly 18,000 entries in a 25-year-old German wiki. The company says misalignment caused "new types of real-world impact" for the first time and plans to release a disclosure framework.

The article OpenAI admits its disclosure practices need work after its autonomous agents hacked a German wiki appeared first on The Decoder.

OpenAI Agents Hacked Another Website

Plus: Tens of millions of US and Canadian drivers’ licenses go up for sale on the dark web, the US military finally tries to tackle the risk online ad data poses to troops, and more.

OpenAI agents discussed ways to escape their sandbox on public wiki

4 September 2026 at 22:17

Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.

In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.

Colluding to share answers

The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.

Read full article

Comments

© Getty Images

Once popular for attacking AI, ASCII smuggling is embraced by spammers

4 September 2026 at 17:18

A clever technique used to hide malicious prompts in attacks on AI agents has been adopted by spammers to evade filters on email platforms that are designed to flag unwanted messages used in mass campaigns.

The technique is broadly known as ASCII smuggling. It gained attention two years ago as a means of making a class of AI attack known as prompt injections more stealthy. Malicious instructions embedded in emails or other untrusted content to be processed by an LLM aren’t written in ordinary text. Instead, they’re rendered by a special range of Unicode tags. For example, the tag point U+E0041 mirrors “A,” and U+E0061 mirrors “a.”

No longer just for obscuring prompt injections

The block of 128 tags mimics a portion of the American Standard Code for Information Interchange almost perfectly, with one major difference: the characters they encode are readable by computers but, by design, are almost completely invisible to humans. By expressing the malicious prompts in these tags, LLMs detect the instructions, but people reading the email never see them. There’s much more about ASCII smuggling here.

Read full article

Comments

© Getty Images

OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits

4 September 2026 at 13:24

According to an analysis by collusion.wiki, autonomous AI agents that identified themselves as OpenAI systems left roughly 18,000 posts in a 25-year-old German wiki between May and July 2026. The agents shared answers, raw data, and a trick that let them break out of their sandbox, built on a faked Microsoft cloud address. A single human moderator deleted dozens of pages every day for weeks, but he couldn't keep up with as many as 400 new entries a day. According to Reuters, OpenAI had known about it for weeks but didn't go public.

The article OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits appeared first on The Decoder.

Prediction Market Betting Is Getting People Banned and Arrested

This week on Uncanny Valley, we dig into the latest prediction market buzz, Flock’s AI-powered police search tool, and how tech bros don’t know how to talk about “rouge” AI agents

Confused about which VPN is right, US senator asks the NSA for guidance

3 September 2026 at 19:52

A prominent US senator is asking the National Security Agency to provide guidance to the general public on best practices for using virtual private networks to secure their communications from spying by foreign adversaries.

VPNs funnel all of a user’s Internet traffic through an encrypted connection to a remote server. The design provides strong assurances that no one between the user and the server can read the encrypted contents. VPNs also allow users to hide their IP addresses from the destination servers they communicate with. While US agencies have previously recommended use of VPNs, none have given recommendations on which ones provide adequate protection.

It's all in the nuances

There are a host of limitations that can undo many of the protections users may think their VPN provides them. For instance, the encrypted tunnel often terminates once a single server decrypts the traffic and sends it on to its final destination. That means the decrypted traffic or the sending and destination IP addresses may be available for snooping by rogue employees or attackers who hack the server. VPNs also don’t encrypt certain types of metadata, such as time stamps, allowing nation-states to build profiles that can be useful in intelligence gathering.

Read full article

Comments

© Getty Images

China-linked hackers backdoored executives' laptops via USB, exploiting a fix companies had but weren't using

A Chinese state-linked hacking group compromised executive laptops at an agricultural industry conference on Hainan Island this spring — not through phishing or a network breach, but by breaking into hotel rooms and booting the machines from a USB stick while the executives were at dinner.

CrowdStrike, which tracks the group as OVERCAST PANDA, disclosed the campaign in its 2026 Threat Hunting Report and detailed the operation's timeline in an interview with VentureBeat at Fal.Con 2026: an intruder entered one room at around 8 p.m. local time and a second room by 9:57 p.m., writing a backdoor called FlowCloud directly to each laptop's storage before rebooting the machines and leaving. There was no network intrusion, no phishing email, and no credential stolen through a login page.

The report dates the intrusions to between March and May 2026, and the timestamps come from Adam Meyers, CrowdStrike's senior vice president of counter adversary operations, who cleared them for publication in the VentureBeat interview. CrowdStrike's OverWatch team disrupted the intrusions and assessed that OVERCAST PANDA will almost certainly continue. FlowCloud itself predates this campaign by years: Proofpoint documented it in 2020, delivered by phishing to U.S. utilities, and NTT Security's SOC has tracked USB-delivered infections at overseas branches of Japanese organizations since early 2022.

Security researchers have called physical-access tampering with an unattended laptop an "evil maid attack," since Joanna Rutkowska demonstrated one with a bootable USB stick in 2009. Physical-access operations are rare across the 290 named adversaries CrowdStrike tracks, according to Meyers, and MUSTANG PANDA's version depends on a dropped USB stick the victim plugs in. What Meyers identified as novel is the combination of hotel-room entry by a state intelligence service with malware deployment, booting the target machine from the USB rather than relying on a user to execute a file from it.

When the executives powered on the next morning, the trigger fired and FlowCloud loaded. Keylogging, screen capture, file collection, and credential harvesting began.

"We have the visibility once the machine boots up," Meyers told VentureBeat. A registry key or similar trigger starts FlowCloud sometime after the operating system loads, and that's when Falcon's sensor picks it up. The gap is the window between the USB write and the next boot — the hours the laptop sits compromised and undetected before the executive logs back in.

CrowdStrike published that gap a month before it announced its AI security product slate — Falcon Guardian, SafeMind, the Agentic Identity Provider, and AI Gateway — at Fal.Con 2026 this week.

Why existing security tools missed it

EDR needs the operating system loaded and the agent running. MFA waits for a login attempt, phishing training for an email, AI agent security for an agent to secure.

OVERCAST PANDA bypassed all of them at the point of entry. The initial compromise completed below the running OS, below the EDR agent, below the authentication stack. Falcon caught FlowCloud once its process started after boot, but by then the implant and its trigger were already on disk.

"Hotel entry is a very common thing," Meyers said. "Talk to any corporate physical security person. They're generally aware of hotel entry, but I think what is unique is the combination of hotel entry with deployment of malware."

Meyers said he thinks China's Ministry of State Security sits behind OVERCAST PANDA. The people entering the rooms are either officers or agents of the MSS or the Ministry of Public Security, or hotel housekeeping staff the services have bribed or compelled, Meyers told VentureBeat. A separate mid-2026 intrusion targeted a U.S.-based media professional using the same tradecraft, according to the report. Targeting an agricultural conference aligns with collection priorities Meyers tied to China's five-year plans.

What CrowdStrike announced at Fal.Con and where runtime security begins

Nvidia CEO Jensen Huang joined George Kurtz on the Fal.Con stage to unveil SafeMind, an agentic cybersecurity system built on Nvidia Nemotron open models and CrowdStrike's threat data. Meyers told the Fal.Con audience that 7,400 CVEs were registered in June 2026, a 96% increase over June 2025, and that CrowdStrike submitted 2,400 of them via responsible disclosure, roughly 30% of all CVEs registered that month.

Falcon Guardian, the company's runtime security layer for AI agents on the endpoint, went live the minute Kurtz put the slide up, CrowdStrike President Mike Sentonas told the Day 2 audience, and AI Gateway, listed as a Guardian capability, ships in September as a hosted service with a hybrid version to follow. AJ Shipley, CrowdStrike's chief product officer, told VentureBeat that CrowdStrike will embed a SafeMind model into Guardian for malicious-prompt detection within the next couple of weeks.

The threats those products address are real, and the report quantifies them. AI agent-triggered detection leads grew at 2.5 times the rate of human-triggered leads, by OverWatch's count. Cloud-conscious eCrime activity surged 171% over the reporting period. Vishing intrusions doubled in the first half of 2026 compared to the second half of 2025, with the eCrime group SNARKY SPIDER moving from account takeover to data exfiltration in under five minutes after compromising SSO-integrated SaaS applications.

Every one of those threats is network-based. All of them assume a running OS, an active user session, or a live cloud workload.

The controls that stop this are firmware and policy

"It's a solvable problem," Meyers said. "It's just an inconvenient solution, which means that a lot of people don't do it."

CrowdStrike itself has shipped firmware attack detection and BIOS settings auditing through the Falcon sensor since May 2019, including a Dell SafeBIOS integration that surfaces BIOS verification telemetry in the Falcon console. The ability to audit security-related BIOS settings on the laptops executives carry has sat inside the platform for seven years. Pointing it at travel devices is a decision, not a product gap.

The controls that would have blunted the OVERCAST PANDA campaign are old and cheap, and each does a different job. Disabling external boot in UEFI removes the vector. A BIOS administrator password keeps it disabled. Pre-boot authentication lets a foreign boot environment load and still keeps the encrypted volume unreadable until a human supplies the PIN or key. Firmware monitoring detects tampering after the fact.

"Don't bring anything with you that you're not comfortable with handing over to a foreign intelligence service," Meyers advised. He used temporary laptops and email accounts on overseas trips while at CrowdStrike, wiping the device when he returned. The exposure starts at customs. Officials can seize a device and compel a login, he added.

"They have master keys to that stuff," was his verdict on hotel safes.

Why scale wins the priority fight

Intrusions tracked by CrowdStrike OverWatch grew about 4% over the reporting period, after a 27% rise the year before, a plateau CrowdStrike attributed to a shift toward more complex, resource-intensive campaigns. The OVERCAST PANDA hotel room operation is the example.

The network threat worries Meyers more. Asked to weigh OVERCAST PANDA's hotel room campaign against the REVENANT SPIDER case he had shown on the Fal.Con stage, an eCrime group using AI to compromise 17 victims with custom web shells in 48 minutes, he picked REVENANT SPIDER.

"You can't intrude on hotel rooms at scale," he said. "You can't intrude on physical devices at scale. And even then, it's just one device." The person in the room is the target, and the intrusion rarely pivots further, he added. "REVENANT SPIDER, they're moving at that speed and they're using AI across the board, and that's a whole other threat, and I think that's more concerning for the average enterprise."

Network-speed, AI-powered intrusions scale. Physical-access tradecraft does not. Security budgets follow the threat that hits the most machines. The threat that is hardest to detect on one machine gets what is left.

But the executives who attended an agricultural conference in China this spring were the specific targets of a state intelligence service, one that chose the slow, unscalable method precisely because it works where network-based attacks fail.

The conference itself is the threat model

Executives at conferences are the campaign's targets, and runtime security starts only once the machine boots. The vendors filling the Las Vegas show floor this week were selling that same runtime protection to attendees whose own laptops carry the identical gap.

Organizational fracture is the real problem. Falcon Guardian ships to one team, and BIOS configuration on travel laptops belongs to another. The Agentic IdP rolls out under identity governance while the decision about whether executives carry production-access machines to international conferences sits with a different group. And the budget line that funds cloud-threat defense has nothing to do with travel-device policies.

Meyers has lived both sides. "I've talked to companies where they're like, we're having a board meeting in Shanghai, and I'm like, why would you do that?"

What security leaders need to do before the next trip

Audit every executive laptop for USB boot status. If the device can be booted from USB right now, it has the same gap OVERCAST PANDA exploited this spring. The steps below cover Windows laptops, the platform FlowCloud targets.

Enforce full-disk encryption with pre-boot authentication. BitLocker in a TPM-only configuration is a documented weak point against physical access. SCRT researchers pulled the volume master key off the LPC bus with a $49 FPGA module in 2021, and Dolos Group did the same over SPI that year. OVERCAST PANDA wrote a backdoor and its post-boot trigger to the Windows volume, so the operators had write access to it. That points to machines that were either unencrypted or protected by a configuration the operators defeated. Pre-boot authentication with a PIN or USB key forces a human step before storage becomes readable.

Verify Secure Boot is enabled and the revocation list is current. Secure Boot validates signatures on boot components and blocks most unauthorized bootloaders, but it leaves external media bootable and signed shims can still carry a bypass. ESET published findings on 11 legacy Microsoft-signed UEFI shims in July 2026 that let untrusted code run at boot on any machine trusting Microsoft's third-party certificate. Microsoft revoked them in its June 9, 2026 DBX update, so a laptop that skipped that update still trusts them. Lock the boot order at the UEFI level, disable one-time boot menus, and set a BIOS administrator password that covers both the setup utility and any boot-override key. Meyers' read is that a lot of these settings go unchecked because the fix is inconvenient.

Issue travel-only devices for international conferences with no access to production systems, no saved credentials for internal tools, and no persistent VPN configuration.

"If they can get their hands on it, they can own it," Meyers put it, citing an old DEF CON adage. Falcon catches FlowCloud only after boot — the exposure is the hours between the USB write and the next login, while the laptop sits closed and compromised.

"It's cheap to buy a couple of laptops and a couple of phones," Meyers said. The controls that close that window are a handful of firmware settings and a spare laptop. The question is whether anyone has deployed them.

Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities

Google keeps cranking out Flash models: the company on Wednesday announced two versions of a new 3.8 Flash

The variants include a standard Flash, a “workhorse” model for agentic tasks, software development, and multi-step reasoning, and Flash Cyber optimized for vulnerability detection and mitigation.

Google CEO Sundar Pichai said in an X post that 3.8 Flash delivers “significant leaps” from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. For instance, it outperformed many large frontier models on the DeepSWE coding benchmark, at far lower cost. 

Meanwhile, Flash Cyber is the company’s “most capable” cybersecurity model, Pichai said; it also matches frontier-level performance when it comes to discovering vulnerabilities and patching them at scale. The model achieved 86.2% on the CyberGym cybersecurity benchmark and 47.2% on CWE-Bench, which evaluates AI patching abilities. In an internal Google benchmark, the model achieved a more than 70% success rate discovering vulnerabilities across 20 programming languages, Pichai said. 

3.8 is Google’s third Flash release in six weeks and comes quickly on the heels of version 3.7. 

3.8 working "harder" with "greater diligence"

3.8 Flash is available now in Gemini Enterprise; devs can try it out in the Gemini API via Google AI Studio, Google Antigravity, Android Studio, or generate UIs in Stitch. It is priced at $0.75 per million input tokens and $3.75 per million output tokens — the same introductory pricing as Gemini 3.7 Flash — and users can customize and adjust model effort levels based on their needs around quality, cost, and latency.

For instance, when compute efficiency is a priority, they can adjust to lower token overhead, or simply continue working with 3.7 Flash, which is “fully supported for efficiency-first workloads,” Google senior product director Tulsee Doshi and Gemini security lead Raluca Ada Popa wrote in a blog post

“3.8 Flash works harder,” exhibiting “greater diligence” with complex tasks like executing extra reasoning steps, although at times it may use more tokens to maximize performance, Doshi and Popa note. The model has a 1M-token input window and a 64K-token output limit, and can ingest text as well as images, audio, video, and PDF files. 

3.8 Flash was evaluated across numerous benchmarks testing coding, multimodal capabilities, computer use, long-context and knowledge work, and scientific reasoning. Google says it also does well in specialized knowledge domains requiring more in-depth analysis and reporting. For instance, the model outperformed its predecessor and other frontier models on benchmarks like Vals Finance Agent V2 for finance, and Harvey's Legal Agent Benchmark for law; it also scored 54.9% on Humanity’s Last Exam (HLE)-Verified, reflecting its ability to take on multi-step reasoning tasks across subjects like math, science, and humanities. 

In one example shared by Google, Gemini 3.8 Flash built a game with a simple prompt using looping techniques in Google’s Antigravity platform. The game uses puzzles, storytelling that changes based on the environment, and images and textures from Nano Banana to create a 3D experience (in this case a wizard navigating a castle). 

In other instances, the model created a fully-functional DOS version of Google Maps featuring interactive locations, directions, and street views; a 3D visualizer that automatically decomposed devices into layers for inspection with a slider capability; and a topographic map of famous geographical sites based on real datasets from the U.S. Geological Survey, complete with real-time cross-sections, 2D projections, and scientific explanations. 

According to Arena.ai, 3.8 Flash landed at No. 14 in Agent Arena, ranking above DeepSeek-V4-Pro, and showed a significant jump over Gemini 3.7 Flash (which sits all the way down at No. 32). It debuted at No. 7 in Text Arena, ahead of Claude Opus 5 and Gemini 3.7 Flash. It improved over 3.7 Flash in several areas: multi-turn requests, writing, literature, and language, longer queries, hard prompts, coding, instruction following, software and IT services, and business, management and financial ops. 

Flash Cyber is already securing Google's code

Flash Cyber is initially being rolled out to “trusted defenders” through Google’s Fairwind Program, which prioritizes government authorities, critical-infrastructure operators, and other partners looking for advanced cyber defense capabilities. Organizations can apply for access. 

Google says the model version has undergone “rigorous training” in the cybersecurity domain and represents a “significant leap in prompt injection robustness.” It is particularly adept at autonomous vulnerability discovery — at least, based on internal Gemini benchmarks — and automated patching. It is also very good at coding, Popa said in a video. 

The goal was to equip defenders with expert-level capabilities to give them a leg up over threat actors (whether malicious, fellow AI agents, or human hackers). “We have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation,” Doshi and Popa explain. 

The model ships a more permissive set of mitigations for cybersecurity safeguards — which is why, for now, it is only being shared with limited partners — and safeguards against misuse in cyber offense and areas like chemical, biological, radiological, and nuclear (CBRN).

Google is already using 3.8 Flash Cyber to secure its own code; it produced 2.6 times more correct patches in Chrome vulnerabilities versus much larger commercial models. 

Wiz — which Google acquired earlier this year at a historic $32 billion — reported that 3.8 Flash Cyber had 7.5% to 9.7% higher recall of real-world vulnerabilities on an internal penetration testing benchmark at 2.3 to 5.2 times lower cost than leading frontier models. Similarly, Google’s Cloud Vulnerability Research found a critical foundational vulnerability in less than 2 hours with 3.8 Flash Cyber. Typically, that research and discovery would take months, Google claims. 

AI agents are “incredibly skilled” at finding and exploiting vulnerabilities, Popa said. Scanning large codebases with big AI models is expensive, and defenders are overwhelmed. “In cybersecurity, attackers need only find one significant flaw over millions of lines of code. Defenders have to remove every one of those flaws to be able to defend against attackers.” 

Doug Turner, engineering director for Chrome, described a “vulnerability apocalypse” in recent months due to generative AI. “Simply overnight, we saw a hockey stick increase in the number of software vulnerabilities reported through our vulnerability research program,” he said in a video. 

One interesting vulnerability 3.8 Flash Cyber discovered had been in Chromium and Chrome for 13 years, he explained. It was a “very subtle bug” that dozens, if not hundreds, of engineers looked at but never flagged. “Gemini 3.8 Flash Cyber is going to allow us to create better suggested fixes so that developers’ lives can get a lot easier.” 

I rented a car, and within hours, my driver's license was for sale

2 September 2026 at 20:32

Not long ago, I rented an SUV from a well-known car rental company. Within hours of an employee scanning my driver's license, a high-resolution scan of my ID was available for sale on the dark web.

An exposé published Tuesday by KrebsOnSecurity reports that my license was one of more than 153 million that were available through Nexus, the name of the new ID theft service. Like other driver's licenses available there—including some belonging to journalist Brian Krebs, his mother, an FBI assistant director, and several security researchers—my license was purported to include multiple image files showing both the front and back of the ID. Besides a basic image scan, the files also captured the images in the infrared and ultraviolet spectrums. Presumably, the additional formats may allow cloned-based counterfeit IDs to pass hologram tests.

Growing by the day

Besides advertising the availability of driver's licenses, Nexus offered to sell a bevy of other forms of ID. They included:

Read full article

Comments

© Nexus

Stolen Claude session cookies can reach corporate Gmail through grants no IT admin can revoke

Infostealers replayed stolen Claude session cookies into paid accounts without ever touching the login page two-factor authentication guards.

The accounts Anthropic flagged were card-billed, self-serve accounts, which is the population no corporate identity provider governs, and no admin console can sign out. Session-cookie replay bypasses SSO as thoroughly as it bypasses 2FA. What SSO provides here is revocation and visibility, not prevention. The company disclosed the campaign in notification emails to affected users, named six stealer families, signed the accounts out, stripped the saved payment methods, and refunded the charges it found.

The burned usage is the small loss. What those sessions could reach is the exposure, and none of it sat behind an identity controlled by an enterprise.

Anthropic told affected users that a bad actor was using common infostealer malware to lift Claude login sessions off their computers and then replaying them to burn the accounts' usage, according to the notification an affected user posted to Reddit and BleepingComputer reported on August 30.

It named Vidar, LummaC2, StealC, RedLine and Acreed on Windows and Atomic Stealer on a small number of Macs, and it described general-purpose malware that copies browser login cookies along with saved passwords. "Your Claude session was likely one of the many things it collected," the email said.

A session cookie is the proof that a login already happened

The attack chain runs in one direction, from an infected machine through a stolen cookie past a checkpoint that never fires, and into everything the account can reach.

Signing the accounts out worked because a replayed cookie dies with the session it copies.

Two-factor authentication guards the login page. The site then hands the browser a cookie so the user stays signed in, and an attacker who copies that cookie and replays it looks to the server like the person who already passed the check. Help Net Security described the mechanism on August 31 as session theft becoming the new credential theft.

Anthropic spotted the theft in the usage meter. Limits were refilled and drained while the owner was away from Claude, the company wrote.

One Redditor who received the notification traced the infection to a pirated game, per BleepingComputer. That is one machine, and Anthropic has not said what the others ran.

Anthropic's notification gave no count. The company had not responded by publication to VentureBeat's questions on how many accounts were affected, whether any Team or Enterprise seats behind SSO were among them, or whether the replayed sessions reached conversation history or connected apps rather than usage alone.

Removing a saved card and refunding charges point to directly billed, self-serve accounts that authenticate through Anthropic's own login rather than a corporate identity provider. Those include personal subscriptions. Team and self-serve Enterprise organizations can also be card-billed, so the deduction is strong rather than closed.

Bugcrowd CEO Dave Gerry told Axios in early August that his company sent employees nearly a dozen emails saying the OpenClaw agent was not allowed on corporate networks, and employees kept trying to download it anyway. A personal Claude subscription on a managed laptop is the same reflex, and it comes with a card on file. LayerX data in Akamai's enterprise AI risk report found 47% of enterprise AI conversations run through personal identities, with Claude at 61%.

The pirated game is one vector. In July, attackers hosted a spoofed Claude download page on the claude.ai domain itself through a public Artifact, and a sponsored Bing ad sent employees searching for "Claude Desktop app" straight to it. Huntress documented the campaign, named FakeAgent, after SectopRAT compromised employees at 29 organizations in two days. The artifact collected roughly 7,100 downloads before Anthropic removed it. A separate campaign pushed a fake Claude installer through a spoofed download site earlier in the year, per Malwarebytes. The vector is not piracy. It is enterprise employees searching for the official app on their work machines.

Refunds cover the usage. Nothing covers the connectors

A replayed session inherits everything the legitimate one could reach, and Anthropic has not said whether these did. On a Claude account, that means the conversation history, the files uploaded into projects, and any connectors the owner authorized. Anthropic's help center states that connectors let Claude retrieve data and take actions inside connected services and that Claude inherits each person's permissions from the connected service. Read and search operations run without approval. Write actions, including send, reply, forward, share, move, and trash, are approval-gated by default. The exfiltration path is the one that is open. Google Workspace connectors are available to individual Claude accounts, so a personal Pro subscription can hold a live authorization into a Gmail inbox or a Drive folder.

If that inbox is the work inbox, the attacker holding the replayed cookie has a read path into it that the corporate identity provider evaluated once, at the moment the employee clicked allow, and rarely again. On a personal plan, the employee owns that grant. No Claude tenant administrator can sign that account out, and the Workspace or Entra administrator who can pull the underlying grant rarely knows it exists.

Adam Meyers, CrowdStrike's senior vice president of counter adversary operations, put numbers to the market in an August 6 Axios interview. Criminals have been buying and reselling stolen ChatGPT, Claude and Gemini credentials since ChatGPT took off in late 2022, fed by infostealer malware. CrowdStrike's 2026 Threat Hunting Report documents one LLMjacking campaign that pushed nearly 200,000 API requests through a compromised cloud account's AI model access in two minutes.

Meyers drew the line in a July briefing on the report. LLMjacking, in his framing, is stealing the credentials, and cost harvesting is what the buyer does next, manipulating AI resources that belong to the victim "in order to conduct operations and generate massive bills as a byproduct of that," he said. "So think of this as LLM coin mining."

One architect refused to build the same exposure into his product

Tom Kleinpeter, co-founder and chief architect at Common Room, described in written answers to VentureBeat why he held his company's AI agent integrations back through the summer of 2025.

"We rejected local MCP servers early, full stop. That path meant storing a long-lived API key or token on someone's machine. Steal that credential, and you can impersonate the user, pull their data, or do anything else the token allows, indefinitely, until someone notices and manually revokes it. We weren't willing to ship that."

Common Room shipped its first agent integration in October 2025 with Okta's Auth0 handling authentication, separate read and write scopes, and writes off by default, per Kleinpeter.

An AI coding agent working on Common Room's own system proposed caching access tokens in plain text in Redis to cut down on repeated authentication calls, he wrote. It worked, and it would have parked live credentials in shared infrastructure had a human reviewer not caught it before it shipped.

Asked what was acceptable in 2024 and a liability now, he named one thing. "Long-lived, broadly scoped API keys. Those made sense when one human operated one trusted system and stayed in the loop. Agents now run across laptops and multiple clients, often with no human watching in real time."

Okta gave agents governed identities the same week Claude users lost their cookies

Okta made Agent SSO generally available on August 24, registering AI agents as first-class identities in Universal Directory and issuing short-lived, identity-governed tokens in place of stored credentials, according to the company's announcement. The release names Claude as its example of an agent a security team can now govern natively.

Six days later, Anthropic was signing users out because six stealer families had copied the humans' Claude cookies. The agents got governed identities. The people using Claude on their own cards did not.

VentureBeat's July Pulse Research wave on agent security found 63% of 116 enterprises report credential sharing somewhere among their AI agents, and 3% run Okta for AI Agents.

That 3% has a reason, Kayne McGladrey, author of the forthcoming "Cyber Risk is a Myth" and a senior IEEE member, told VentureBeat during a July interview. "It's only those well-resourced companies that are above the poverty line that have met all the prerequisites," he said.

The prerequisites he named are the same controls most enterprises still treat as hygiene, not strategic investment.

"If they don't have their defenses in order, like attack surface management or blast radius containment or basic MFA, that would not be a useful capability or a meaningful spend."

Anthropic's position deserves its hearing. The company told users it has no reason to believe the malware is related to Claude, installed through Claude, or tied to anything they did with Claude, and it warned that signing out stops the stolen sessions while leaving the malware in place to steal the next login.

Both hold, and they are the last thing a provider can do, because the infected device belongs to the customer. On a work laptop, the device belongs to the enterprise, and the control that catches Vidar or LummaC2 before it reads a cookie jar is endpoint detection, the control in this story the security team already runs.

The profession's gap is rarely a missing control anymore, in McGladrey's framing. "I think we've got technical solutions for nearly all of the things that could go wrong, what we don't have is a way of prioritizing those," he argued.

The endpoint team owns the machine. The identity team owns an SSO the account never touched, and the AI governance lead wrote a policy the employee routed around the day the card went on file.

Each of those owners is paid to close a different gap. "Engineering is comped on getting product out the door quickly, your internal audit team is comped on checking boxes to meet your compliance goals, and security is comped and sometimes penalized on a lack of incidents," he argued. "People aren't doing the wrong thing either. They're doing what pays their bills on an ongoing basis."

What security leaders need to do next

Add AI accounts to the infostealer response playbook. When an endpoint alert names a stealer family, treat every AI service session on that machine as compromised, revoke what the enterprise tenant lets you revoke, and have the employee sign out of personal accounts until the machine is clean.

Warn users that the notification itself is now a phishing template. Help Net Security flagged copycat phishing impersonating Anthropic using this campaign as pretext. If the notification lands in a user's inbox, the next email that looks like it may not be from Anthropic.

Count the personal subscriptions on managed devices. Browser telemetry, CASB logs, and expense reports surface the sessions and the payments.

Stop personal AI accounts from holding OAuth grants into corporate Google Workspace or Microsoft 365. Both platforms let administrators restrict third-party app authorization. Use that gate so a work inbox can only be attached from a tenant the security team can revoke.

Revoke the OAuth grants Claude already holds, not just the Claude session. Signing out of Claude invalidates the stolen session but does not revoke the Google or Microsoft grant Claude was already authorized to use. Check Google's third-party app authorizations and Microsoft's enterprise application consents for live grants the sign-out left behind.

Move the heavy users onto the organization-managed tenant. On Team and Enterprise plans, an owner decides whether connectors can be enabled at all.

Put session binding on the renewal agenda. Google shipped Device Bound Session Credentials in Chrome 146 on Windows in April and turned it on by default for Google accounts and Workspace Individual accounts in May, binding each session to a private key in the device's TPM so a copied cookie cannot be refreshed anywhere else. It covers Chrome on Windows only so far, so the Mac victims in this campaign sit outside it. Ask Anthropic and OpenAI for parity and Google for a coverage date before the next contract signs.

Anthropic sent its notification to individuals. The laptop the cookie came from belongs to whoever manages it, and Vidar and LummaC2 will be back for the next login on the same machine.

BGP hijack infecting networks caused by a comedy of errors that’s not funny at all

2 September 2026 at 11:00

Hackers carried out a supply chain attack that installed malware on networks using an unusual technique: hijacking a chunk of Internet space where cloud management software used by hosting providers, data centers, and other large infrastructure companies is updated.

In a well-coordinated operation, the unknown attackers exploited weaknesses in the routing security setup of hosting provider Hetzner Online and the process for attaining valid TLS certificates. The lapses allowed the attackers to successfully perform a BGP (Border Gateway Protocol) hijacking to obtain control over IP addresses assigned to Softaculous. The company, based in the United Arab Emirates, is the maker of a platform for installing and managing Web software and is the developer of Virtualizor, a management platform for virtualized environments.

Softaculous used the IPs to issue updates and host a client and billing site. With control over the hijacked space, the attacker was now using the addresses to push malware masquerading as updates to unsuspecting users.

Read full article

Comments

© Getty Images

❌