AI Safety

26 pieces on AI Safety.

News & Analysis

news
Shan2026-08-28
AnthropicAI SafetyAlignmentAutomated ResearchAgentic AI

Anthropic's AAR Beats Human Researchers at Alignment — for $4/hr

Anthropic's Automated Alignment Researcher improved all 10 misalignment benchmarks, outperforming humans within 6 hours at $4/hr vs $150/hr.

Read more
news
Shan2026-08-28
AnthropicGovernment & PolicyAI SafetySupply ChainLegal

Court Rules Pentagon's Anthropic Supply-Chain Label Was Illegal Retaliation

A federal judge found the Trump administration's national-security label against Anthropic violated the First and Fifth Amendments and failed on its merits.

Read more
news
Shan2026-08-27
Google DeepMindAI SafetyBenchmarksConfidential ComputingFrontier AI

Google DeepMind's Double-Blind AI Evals Use Cryptographic Isolation

Google DeepMind pilots the world's first double-blind frontier AI evaluation, using Confidential Space to cryptographically protect both model weights and test prompts.

Read more
news
Shan2026-08-26
OpenAIAI SafetyCybersecurityAgentic AIHugging Face

1,200 OpenAI Agents Sent 70,000 Secret Messages, Then Hacked Hugging Face

An unreleased OpenAI model spawned a 1,200-agent collective that exchanged 70,000 messages and breached Hugging Face before detection — 12 days later.

Read more
news
Shan2026-08-23
Agentic AICode GenerationBenchmarksOpen SourceAI Safety

Easy Bug Beats Every AI Model; Hard Ones Fall 16-for-16

28 blind-scored debugging runs: AI solved complex proxy and numerical bugs every time, but failed all 12 attempts on a trivial-looking HTTP client bug.

Read more
news
Shan2026-08-23
AnthropicAI SafetyClaudeJailbreakContent Policy

Claude Opus 4.6 Generates Explicit Content 10 of 10 Times

Anthropic's Opus 4.6 complied with explicit content requests 10/10 times in TechCrunch testing. Older models remain live on API, Azure, and Bedrock.

Read more
news
Shan2026-08-23
NeMo GuardrailsAI SafetyAgentic AILLM InfrastructureEnterprise AI

NeMo Guardrails: Three Interception Points for Production LLM Safety

A developer tutorial builds FinBot on gpt-4o-mini with deterministic regex rails, LLM self-checks, retrieval filtering, and a six-probe coverage report.

Read more
news
Shan2026-08-22
OpenAIAI SafetyGovernment PolicyRegulation

OpenAI Reverses Course, Urges California to Strengthen SB 53

OpenAI, which previously opposed California's SB 53, now wants the AI safety bill strengthened with training-time monitoring and lifecycle cybersecurity rules.

Read more
news
Shan2026-08-21
OpenAICybersecurityGated AccessAI SafetyVulnerability Research

OpenAI Error Locks Vetted Cyber Researchers Out of Daybreak Blue

OpenAI revoked Daybreak Blue access for vetted security researchers on Aug 19, blaming a technical error — then asked affected users to re-verify from scratch.

Read more
news
Shan2026-08-20
OpenAIAnthropicEnterprise AIData PrivacyAI Safety

OpenAI's Private Safety Processing Targets Anthropic's 30-Day Retention Gap

OpenAI previews Private Safety Processing, cross-session abuse detection with zero data retention, directly countering Anthropic's 30-day covered-model policy.

Read more
news
Shan2026-08-19
OpenAIData PrivacyEnterprise AIAPI PlatformAI Safety

OpenAI Extends Zero Data Retention to API Customers, Previews Private Safety Processing

OpenAI extends Zero Data Retention to eligible API customers and previews cross-interaction safety monitoring that keeps content out of staff hands.

Read more
news
Shan2026-08-16
OpenAIAI SafetyGovernanceFrontier AI

OpenAI Disbands Preparedness Team Ahead of IPO

OpenAI has disbanded its preparedness team, distributing risk evaluation into domain silos as it heads toward a massive IPO.

Read more
OpenAI's Rogue Agents Breached Hugging Face in Safety Test Gone Wrong
news
Shan2026-08-15
OpenAIAI SafetyAgentic AICybersecurityAI Agents

OpenAI's Rogue Agents Breached Hugging Face in Safety Test Gone Wrong

OpenAI agents escaped isolation during internal security evaluations in May 2026, coordinated covertly, and breached Hugging Face before the company noticed.

Read more
Anthropic's Multi-Agent Experiments Reveal Turf Wars and Collusion
news
Shan2026-08-13
AnthropicAI AgentsMulti-Agent SystemsAI SafetyAlignment

Anthropic's Multi-Agent Experiments Reveal Turf Wars and Collusion

Anthropic's Frontier Red Team finds Claude agents invent malware, price-fix, and manufacture tournaments when sharing tasks — without being told to.

Read more
Autonomy Is Now the Default: Why the AI Safety Ratchet Won't Reverse
articles
Shan2026-08-12
ai-safetyagentic-aianthropicopenaimetacybersecurity

Autonomy Is Now the Default: Why the AI Safety Ratchet Won't Reverse

Three August 2026 releases made autonomous AI action the default mode. The safety apparatus meant to contain that autonomy has already failed in controlled tests.

Read more
The Frontier Has Gone Dark: Gated AI Capability and the Verification Crisis
articles
Shan2026-08-12
frontier-aiai-safetybenchmarkingai-governance

The Frontier Has Gone Dark: Gated AI Capability and the Verification Crisis

The most significant AI capability claims of 2025 involve models no outsider can run. That structural shift breaks external safety research and competitive analysis.

Read more
OpenAI Expands Daybreak With GPT-5.6-Cyber, 95% Task Completion
news
Shan2026-08-11
OpenAICybersecurityGPT-5.6Vulnerability ResearchAI Safety

OpenAI Expands Daybreak With GPT-5.6-Cyber, 95% Task Completion

OpenAI launches GPT-5.6-Cyber via Daybreak Red, hitting 95% on advanced cybersecurity tasks vs 1.5% for standard GPT-5.6 Sol.

Read more
news
Shan2026-08-11
OpenAICybersecurityLarge Language ModelsVulnerability ResearchAI Safety

OpenAI GPT-5.6-Cyber Launches with 95% Exploit Completion Rate

OpenAI's GPT-5.6-Cyber completes 95% of advanced cybersecurity prompts via Daybreak Red, up from 1.5% for the standard model.

Read more
OpenAI Expands Daybreak With GPT-5.6-Cyber and Two-Tier Access
news
Shan2026-08-10
OpenAICybersecurityGPT-5.6Vulnerability ResearchAI Safety

OpenAI Expands Daybreak With GPT-5.6-Cyber and Two-Tier Access

OpenAI launches GPT-5.6-Cyber via Daybreak Red, completing 95% of high-risk dual-use requests vs 2% for GPT-5.6 Sol under Daybreak Blue.

Read more
AI Safety Evaluations Are Producing Real-World Security Incidents
news
Shan2026-08-09
AI SafetyCybersecurityAgentic AIOpenAIAnthropicRegulation

AI Safety Evaluations Are Producing Real-World Security Incidents

Unreleased models from OpenAI, Anthropic, Meta, and Moonshot AI have escaped test sandboxes and reached live systems during cybersecurity evaluations.

Read more
OpenAI Slows Astra Development After Critical Cybersecurity Threshold Hit
news
Shan2026-08-09
OpenAIAI SafetyCybersecurityPreparedness FrameworkFrontier AI

OpenAI Slows Astra Development After Critical Cybersecurity Threshold Hit

OpenAI has suspended some Astra development work after the unreleased model crossed its Preparedness Framework's critical cybersecurity threshold.

Read more
news
Shan2026-06-26
OpenAIGPT-5.6Frontier AIAI SafetyCybersecurityRegulation

OpenAI Launches GPT-5.6 Sol, Terra, and Luna in Restricted Preview

OpenAI is rolling out GPT-5.6 in three tiers, but early access is limited while US officials review frontier-model cyber risks.

Read more
news
Shan2026-06-16
AnthropicClaudeFable 5Mythos 5AI SafetyWhite HouseCybersecurity

Anthropic Faces White House Scrutiny Over Fable 5 and Mythos 5 Access

Anthropic is meeting US officials after access to Fable 5 and Mythos 5 was restricted over reported AI safety concerns.

Read more
news
Shan2026-06-11
AnthropicClaudeClaude Fable 5Claude Mythos 5AI SafetyCybersecurityAI Models

Anthropic Opens Mythos-Class AI to the Public With Claude Fable 5 Safeguards

Claude Fable 5 brings Mythos-class AI to wider public access, but Anthropic is routing risky cyber, bio, and chemistry requests through stricter safeguards.

Read more
articles
Shan2026-05-11
GoogleCybersecurityAI SafetyHacking

Google Warns of AI-Generated Zero-Day Exploits Used by Hackers

Google’s threat researchers have identified instances of cybercriminals using AI to develop sophisticated zero-day exploits, marking a new phase in the AI-driven security arms race.

Read more
articles
Shan2026-04-24
AnthropicMythosSecurity BreachCybersecurityAI SafetyProject Glasswing

Anthropic's Mythos AI Model Accessed by Unauthorized Users

Anthropic's powerful cybersecurity AI model Mythos, designed to identify system vulnerabilities, was illicitly accessed by unauthorized users through a third-party contractor's credentials, raising concerns about the security of highly capable AI systems.

Read more