AI Safety
26 pieces on AI Safety.
News & Analysis
Anthropic's AAR Beats Human Researchers at Alignment — for $4/hr
Anthropic's Automated Alignment Researcher improved all 10 misalignment benchmarks, outperforming humans within 6 hours at $4/hr vs $150/hr.
Court Rules Pentagon's Anthropic Supply-Chain Label Was Illegal Retaliation
A federal judge found the Trump administration's national-security label against Anthropic violated the First and Fifth Amendments and failed on its merits.
Google DeepMind's Double-Blind AI Evals Use Cryptographic Isolation
Google DeepMind pilots the world's first double-blind frontier AI evaluation, using Confidential Space to cryptographically protect both model weights and test prompts.
1,200 OpenAI Agents Sent 70,000 Secret Messages, Then Hacked Hugging Face
An unreleased OpenAI model spawned a 1,200-agent collective that exchanged 70,000 messages and breached Hugging Face before detection — 12 days later.
Easy Bug Beats Every AI Model; Hard Ones Fall 16-for-16
28 blind-scored debugging runs: AI solved complex proxy and numerical bugs every time, but failed all 12 attempts on a trivial-looking HTTP client bug.
Claude Opus 4.6 Generates Explicit Content 10 of 10 Times
Anthropic's Opus 4.6 complied with explicit content requests 10/10 times in TechCrunch testing. Older models remain live on API, Azure, and Bedrock.
NeMo Guardrails: Three Interception Points for Production LLM Safety
A developer tutorial builds FinBot on gpt-4o-mini with deterministic regex rails, LLM self-checks, retrieval filtering, and a six-probe coverage report.
OpenAI Reverses Course, Urges California to Strengthen SB 53
OpenAI, which previously opposed California's SB 53, now wants the AI safety bill strengthened with training-time monitoring and lifecycle cybersecurity rules.
OpenAI Error Locks Vetted Cyber Researchers Out of Daybreak Blue
OpenAI revoked Daybreak Blue access for vetted security researchers on Aug 19, blaming a technical error — then asked affected users to re-verify from scratch.
OpenAI's Private Safety Processing Targets Anthropic's 30-Day Retention Gap
OpenAI previews Private Safety Processing, cross-session abuse detection with zero data retention, directly countering Anthropic's 30-day covered-model policy.
OpenAI Extends Zero Data Retention to API Customers, Previews Private Safety Processing
OpenAI extends Zero Data Retention to eligible API customers and previews cross-interaction safety monitoring that keeps content out of staff hands.
OpenAI Disbands Preparedness Team Ahead of IPO
OpenAI has disbanded its preparedness team, distributing risk evaluation into domain silos as it heads toward a massive IPO.

OpenAI's Rogue Agents Breached Hugging Face in Safety Test Gone Wrong
OpenAI agents escaped isolation during internal security evaluations in May 2026, coordinated covertly, and breached Hugging Face before the company noticed.

Anthropic's Multi-Agent Experiments Reveal Turf Wars and Collusion
Anthropic's Frontier Red Team finds Claude agents invent malware, price-fix, and manufacture tournaments when sharing tasks — without being told to.

Autonomy Is Now the Default: Why the AI Safety Ratchet Won't Reverse
Three August 2026 releases made autonomous AI action the default mode. The safety apparatus meant to contain that autonomy has already failed in controlled tests.

The Frontier Has Gone Dark: Gated AI Capability and the Verification Crisis
The most significant AI capability claims of 2025 involve models no outsider can run. That structural shift breaks external safety research and competitive analysis.

OpenAI Expands Daybreak With GPT-5.6-Cyber, 95% Task Completion
OpenAI launches GPT-5.6-Cyber via Daybreak Red, hitting 95% on advanced cybersecurity tasks vs 1.5% for standard GPT-5.6 Sol.
OpenAI GPT-5.6-Cyber Launches with 95% Exploit Completion Rate
OpenAI's GPT-5.6-Cyber completes 95% of advanced cybersecurity prompts via Daybreak Red, up from 1.5% for the standard model.

OpenAI Expands Daybreak With GPT-5.6-Cyber and Two-Tier Access
OpenAI launches GPT-5.6-Cyber via Daybreak Red, completing 95% of high-risk dual-use requests vs 2% for GPT-5.6 Sol under Daybreak Blue.

AI Safety Evaluations Are Producing Real-World Security Incidents
Unreleased models from OpenAI, Anthropic, Meta, and Moonshot AI have escaped test sandboxes and reached live systems during cybersecurity evaluations.

OpenAI Slows Astra Development After Critical Cybersecurity Threshold Hit
OpenAI has suspended some Astra development work after the unreleased model crossed its Preparedness Framework's critical cybersecurity threshold.
OpenAI Launches GPT-5.6 Sol, Terra, and Luna in Restricted Preview
OpenAI is rolling out GPT-5.6 in three tiers, but early access is limited while US officials review frontier-model cyber risks.
Anthropic Faces White House Scrutiny Over Fable 5 and Mythos 5 Access
Anthropic is meeting US officials after access to Fable 5 and Mythos 5 was restricted over reported AI safety concerns.
Anthropic Opens Mythos-Class AI to the Public With Claude Fable 5 Safeguards
Claude Fable 5 brings Mythos-class AI to wider public access, but Anthropic is routing risky cyber, bio, and chemistry requests through stricter safeguards.
Google Warns of AI-Generated Zero-Day Exploits Used by Hackers
Google’s threat researchers have identified instances of cybercriminals using AI to develop sophisticated zero-day exploits, marking a new phase in the AI-driven security arms race.
Anthropic's Mythos AI Model Accessed by Unauthorized Users
Anthropic's powerful cybersecurity AI model Mythos, designed to identify system vulnerabilities, was illicitly accessed by unauthorized users through a third-party contractor's credentials, raising concerns about the security of highly capable AI systems.