Local Inference
6 pieces on Local Inference.
News & Analysis
FreeToken Runs 284B MoE Models on a Single Consumer GPU
FreeToken's q* policy and semantic anchor checkpointing deliver 3–4× faster decode and 6–30× faster prefill for frontier MoE models on RTX consumer hardware.
Perplexity Portable Computer: Full Agent Harness on DGX Spark, $0 Local Steps
Perplexity ships a full agentic harness on NVIDIA DGX Spark with zero per-token cost for local steps and OS-enforced sandboxing.

Qwen 3.8 27B Is Strong but Overthinks by Default
Alibaba's 17 GB Qwen 3.8 27B excels at vision, tool use, and coding agents — but its xhigh reasoning default burns tokens on trivial prompts.

Meta Releases Muse Glimmer, a 30B Open-Weight Local Agent Model
Meta's Muse Glimmer is a 30B Apache 2.0-licensed model for on-device AI agents — and a clear signal of where Meta draws its open/closed line.

webAI TwIL-LM: 1.7B and 3B Formal-Logic Models for Local Hardware
webAI releases TwIL-LM, a 1.7B LoRA adapter and 3B merged model for autoformalization, running on 4 GB VRAM under a non-commercial license.

Meta Releases Muse Glimmer: 30B Open-Weight Local Agent Model
Meta's Apache 2.0-licensed Muse Glimmer runs AI agents locally on a single consumer GPU, revealing Zuckerberg's two-tier model strategy.