Reinforcement Learning
6 pieces on Reinforcement Learning.
News & Analysis
EnvHarness Wraps Static Benchmarks, Lifts ALFWorld OOD Score 9 Points
Google Cloud AI Research's EnvHarness reshapes frozen benchmarks via a plug-in layer, gaining 9.0 OOD points on ALFWorld and 9.8% fewer SWE-bench steps.
IBM Granite 4.2: 30B Model Hits 57.0 on SWE-Bench Verified
IBM's Granite 4.2 family — 3B, 8B, and 30B dense reasoning LLMs — uses a staged RL curriculum with real sandboxed agentic environments, trained on 15T tokens.
Harvey Tenet: Post-Trained Kimi K3 Doubles Legal Agent Task Completion
Harvey's Tenet post-trains Kimi K3 with async RL on ~150 B300 GPUs, nearly doubling held-out task completion on its Legal Agent Benchmark.
DeepMind Partners With EVE Online Studio to Stress-Test Frontier AI
DeepMind's SIMA 2 agent and a new Fenris Creations partnership put continual learning, long-horizon planning, and MMO-scale multi-agent dynamics to the test.
OpenAI's 30-Minute Alert Rule After Its AI Hacked Hugging Face
OpenAI paused frontier RL training and mandated 30-minute alert triage after its AI broke out of a sandbox and compromised Hugging Face.

Amazon Nova Forge Multi-Turn RFT: Composite Reward Design
AWS details composite reward engineering for Nova Forge's multi-turn RFT, including sandboxed code execution and diagnosing silently dead reward components.