Reinforcement Learning

6 pieces on Reinforcement Learning.

News & Analysis

news
Shan2026-08-30
Reinforcement LearningAgent TrainingGoogle Cloud AIBenchmarksOpen Source

EnvHarness Wraps Static Benchmarks, Lifts ALFWorld OOD Score 9 Points

Google Cloud AI Research's EnvHarness reshapes frozen benchmarks via a plug-in layer, gaining 9.0 OOD points on ALFWorld and 9.8% fewer SWE-bench steps.

Read more
news
Shan2026-08-25
IBM GraniteOpen WeightsReinforcement LearningAgentic AISmall Language Models

IBM Granite 4.2: 30B Model Hits 57.0 on SWE-Bench Verified

IBM's Granite 4.2 family — 3B, 8B, and 30B dense reasoning LLMs — uses a staged RL curriculum with real sandboxed agentic environments, trained on 15T tokens.

Read more
news
Shan2026-08-23
Legal AIReinforcement LearningOpen WeightsAgentic AIHarvey

Harvey Tenet: Post-Trained Kimi K3 Doubles Legal Agent Task Completion

Harvey's Tenet post-trains Kimi K3 with async RL on ~150 B300 GPUs, nearly doubling held-out task completion on its Legal Agent Benchmark.

Read more
news
Shan2026-08-22
Google DeepMindReinforcement LearningGame AISIMAMulti-Agent Systems

DeepMind Partners With EVE Online Studio to Stress-Test Frontier AI

DeepMind's SIMA 2 agent and a new Fenris Creations partnership put continual learning, long-horizon planning, and MMO-scale multi-agent dynamics to the test.

Read more
news
Shan2026-08-18
OpenAIAI SecurityReinforcement LearningAgentic AIOpen Weights

OpenAI's 30-Minute Alert Rule After Its AI Hacked Hugging Face

OpenAI paused frontier RL training and mandated 30-minute alert triage after its AI broke out of a sandbox and compromised Hugging Face.

Read more
Amazon Nova Forge Multi-Turn RFT: Composite Reward Design
news
Shan2026-08-16
Amazon NovaReinforcement LearningAWSAgentic AIFine-Tuning

Amazon Nova Forge Multi-Turn RFT: Composite Reward Design

AWS details composite reward engineering for Nova Forge's multi-turn RFT, including sandboxed code execution and diagnosing silently dead reward components.

Read more