Reading time: 3 minutes
TLDR: Two Anthropic safety researchers resigned within days of each other this week — Jacob Coxon on September 9 and Joe Benton on September 11 — both warning publicly that Anthropic is hiding safety incidents. The same week, a Russian-speaking threat actor used hundreds of AI agents built on OpenAI's Codex and a DeepSeek model to compromise at least 440 organisations across 48 countries. And Google's AI Agent Development Kit had a CVSS 10.0 remote-code execution vulnerability disclosed on September 9. This is the week AI safety accountability became front-page news.
Joe Benton's resignation statement is worth reading directly. Published on September 11 with 2.03 million views and 22,400 likes, it reads: "I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this." That is a safety researcher who spent years working on AI alignment at the company most publicly associated with AI safety, saying publicly that he left because he believes the company is not being transparent about safety incidents.
Jacob Coxon's exit two days earlier made the same point more quietly. Two departures in two days from the same safety team, both with public statements, is not a coincidence. It is a pattern that says something about the internal dynamics at Anthropic at the moment the company is preparing for an IPO and has committed $80 billion in compute in a single week.
Anthropic has not responded publicly in detail. The company's position — that it is the safety-focused AI lab — is now being contested from the inside by its own safety researchers. For knowledge workers who choose their AI tools partly based on which companies they trust, this is signal, not noise.
AI agents were used to compromise 440 organisations
GreyNoise researchers documented a Russian-speaking threat actor who used hundreds of AI agents built on OpenAI's Codex and a DeepSeek model to exploit two CVEs — CVE-2026-81578 and CVE-2026-82078 — compromising at least 440 PaperCut NG/MF instances at 395 organisations across 48 countries. The campaign began in August and continued into September.
The significance is architectural. This is the first publicly documented large-scale attack campaign where AI agents were not a tool used occasionally to assist human attackers — they were the primary execution layer, running autonomously at scale across hundreds of targets simultaneously. The attack surface of AI-assisted cyberattacks is qualitatively different from traditional automated attacks because the agents can adapt, reason about obstacles, and find novel paths that pre-scripted tools cannot.
For organisations using PaperCut: patch immediately. For everyone else: the threat model for network security now includes autonomous AI agents as first-class attackers, not theoretical future threats.
Google's AI dev kit had a perfect CVSS score vulnerability
CVE-2026-79696, disclosed September 9 with a CVSS score of 10.0 — the maximum possible — lets an unauthenticated remote attacker execute arbitrary code against Google Cloud's Agent Development Kit for Python versions 2.0.0 through 2.6.0 via a crafted test session replay. CVSS 10.0 means: no authentication required, network accessible, complete compromise of confidentiality, integrity, and availability.
If you or your organisation uses Google Cloud ADK for Python in any of those versions, update immediately. The vulnerability affects the tooling that developers use to build AI agents, not the agents themselves — but the blast radius of a compromised development environment is the entire production system that environment builds.
The pattern this week describes
Safety researchers resigning and warning of hidden incidents. AI agents used to compromise hundreds of organisations at scale. A perfect CVSS score in AI development tooling. The theme that ties these together is the same one from last week and the week before: AI capability is advancing faster than the accountability infrastructure — internal governance, external oversight, security practices, and regulatory frameworks — designed to contain it. That gap is where the risks concentrate. The knowledge workers and organisations that close it on their own terms, before they are forced to, will be in a better position than those who discover it the hard way.
P.S. If you run a newsletter or are thinking about starting one, the platform behind AI Quiet Signal is beehiiv. It handles the infrastructure so you can focus on the signal.