Author
Shan
Shan edits AI Mastery's daily reporting and evaluates new models, developer tools, and infrastructure through an engineering-first lens.
Shan edits AI Mastery's daily reporting and evaluates new models, developer tools, and infrastructure through an engineering-first lens.
Recent work
DeepSeek-V4.1-Flash: 890 Bytes Per Token, 437x Smaller KV Cache
DeepSeek's 552B MoE model cuts global KV cache to 890 bytes per token — 437x below V1 — using CED, CSA2, and FP4 quantization.
GPT-6 Astra: $10/M Tokens, 57.9% Terminal-Bench, Critical Cyber Flag
OpenAI's GPT-6 Astra scores 57.9% on Terminal-Bench 4.0, costs $10/$50 per million tokens, and is the first model to hit the Critical cybersecurity threshold.
IBM's 385M-Param PatchTST-FM-r2 Ranks 2nd on GIFT-Eval Zero-Shot
IBM's Granite PatchTST-FM-r2 ranks 2nd zero-shot on GIFT-Eval under Apache 2.0/OpenMDW 1.0, beating larger pretrained rivals.
Paul Christiano, RLHF Co-Inventor and AI Doomer, Joins OpenAI Board
The Alignment Research Center founder joins OpenAI's Safety and Security Committee, which holds final authority over model releases like Astra.
Google's Mantis Cuts AI Security False Positives With Sandbox-First Pipeline
Google open-sources Mantis, a slash-command toolkit that grounds AI vulnerability findings in sandbox execution, cutting token overhead by over 85%.
GPT-6 Astra Hits Amazon Bedrock With 1M-Token Context and Critical Security Tier
GPT-6 Astra is now generally available on Amazon Bedrock with a 1M-token context window, chip-level operator isolation, and OpenAI's first Critical cybersecurity classification.
Meta Muse Runs Each User's Agent in a Dedicated Secure Cloud VM
Meta's Muse gives every user an isolated cloud VM, a Sentinel control plane, and surrogate credentials — with a $130,000 bug bounty on injection attacks.
Mistral's €3B Series D Targets 1 GW of European Compute by 2030
Mistral AI closes a €3B Series D at a €21B+ valuation, led by Samsung, targeting 1 GW of European compute capacity by 2030.
150M-Parameter BDH-CQ Scores 29.2% on ARC-AGI-1 at $0.0007 per Task
Pathway's 150M-parameter BDH-CQ model scores 29.2% pass@2 on ARC-AGI-1 at $0.0007 per task, trained on SageMaker HyperPod with H200 GPUs.
2.4T-Parameter Qwen3.8 Runs on One Node With vLLM and NVFP4
AWS documents a single-node serving path for Qwen3.8-2.4T-A95B on a p6-b300 instance using NVFP4 quantization, cutting TTFT by nearly 60%.
SageMaker Feature Store's UpdateRecord Ends Read-Modify-Write Cycles
AWS ships UpdateRecord for SageMaker Feature Store: atomic partial writes on up to 100 features, no GetRecord required, available in all regions today.
Unikraft Fits 1,024,132 Firecracker VMs on One 48-Core Server
Unikraft's unikernel microVMs cold-resume in ~10 ms from snapshots, enabling over 1 million scale-to-zero sandboxes on a single server.