Reasoning Models
5 pieces on Reasoning Models, including 2 step-by-step guides.
Guides
How to Run DeepSeek R1 Locally with Ollama: Full Setup Guide
Install DeepSeek R1 locally using Ollama in under 5 minutes. Covers model variant selection from 1.5B to 671B, visible chain-of-thought reasoning, REST API usage, Python integration, and building a simple RAG application.
How to Run Qwen3 Locally with Ollama: Setup, API, and a Gradio App
Set up Qwen3 locally in minutes using Ollama. Covers every model variant, thinking mode control with /think and /no_think tags, CLI, REST API, Python SDK, and a practical Gradio reasoning app.
News & Analysis
Ember-1 Cuts Kimi K3 Output Tokens 39% at Identical Price
Fireworks AI post-trained Kimi K3 into Ember-1, cutting output tokens 39% in production while matching task scores — at the same $15/M output price.
ThinkingCap-Qwen3.8-27B Cuts Thinking Tokens 37.2% for 0.86pp Accuracy
BottleCap AI's fine-tune of Qwen3.8-27B drops thinking tokens 37.2% across 12 benchmarks, losing just 0.86pp of macro accuracy at xhigh effort.
Claude Fable 5.1 Hits 52.6% on Science Bench; Max Run Costs $3.30
Anthropic's Claude Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1 and ships five reasoning tiers ranging from ~$0.10 to $3.30 per prompt.