Local Inference

6 pieces on Local Inference.

News & Analysis

news
Shan2026-08-29
Model InferenceOpen WeightsConsumer HardwareMixture-of-ExpertsLocal Inference

FreeToken Runs 284B MoE Models on a Single Consumer GPU

FreeToken's q* policy and semantic anchor checkpointing deliver 3–4× faster decode and 6–30× faster prefill for frontier MoE models on RTX consumer hardware.

Read more
news
Shan2026-08-25
PerplexityNVIDIAAgentic AILocal InferenceOpen Weights

Perplexity Portable Computer: Full Agent Harness on DGX Spark, $0 Local Steps

Perplexity ships a full agentic harness on NVIDIA DGX Spark with zero per-token cost for local steps and OS-enforced sandboxing.

Read more
Qwen 3.8 27B Is Strong but Overthinks by Default
news
Shan2026-08-16
Open WeightsLocal InferenceLLM ReasoningQwenCoding Agents

Qwen 3.8 27B Is Strong but Overthinks by Default

Alibaba's 17 GB Qwen 3.8 27B excels at vision, tool use, and coding agents — but its xhigh reasoning default burns tokens on trivial prompts.

Read more
Meta Releases Muse Glimmer, a 30B Open-Weight Local Agent Model
news
Shan2026-08-11
MetaOpen WeightsAI AgentsLocal InferenceLarge Language Models

Meta Releases Muse Glimmer, a 30B Open-Weight Local Agent Model

Meta's Muse Glimmer is a 30B Apache 2.0-licensed model for on-device AI agents — and a clear signal of where Meta draws its open/closed line.

Read more
webAI TwIL-LM: 1.7B and 3B Formal-Logic Models for Local Hardware
news
Shan2026-08-11
Open WeightsSmall Language ModelsFormal ReasoningAutoformalizationLocal Inference

webAI TwIL-LM: 1.7B and 3B Formal-Logic Models for Local Hardware

webAI releases TwIL-LM, a 1.7B LoRA adapter and 3B merged model for autoformalization, running on 4 GB VRAM under a non-commercial license.

Read more
Meta Releases Muse Glimmer: 30B Open-Weight Local Agent Model
news
Shan2026-08-10
MetaOpen WeightsAI AgentsLocal InferenceLarge Language Models

Meta Releases Muse Glimmer: 30B Open-Weight Local Agent Model

Meta's Apache 2.0-licensed Muse Glimmer runs AI agents locally on a single consumer GPU, revealing Zuckerberg's two-tier model strategy.

Read more