local-llm
5 pieces on local-llm, including 5 step-by-step guides.
Guides
Host and Run Qwen3.8-Flash Locally: Complete Unsloth and llama.cpp Deployment Guide
Run Qwen3.8-Flash locally on CPU, unified memory, or GPUs using Unsloth and llama.cpp with Multi-Token Prediction (MTP) for up to 1.7x faster inference.
Build a Local LLM Zero-Shot Classifier You Can Actually Deploy
Learn how to run zero-shot text classification on a local model with Ollama, enforce strict JSON outputs, and add confidence-aware routing for production triage.
The Best Small Language Models Available on Hugging Face
A comprehensive breakdown of the most capable sub-8B parameter local language models currently dominating the Hugging Face hub.
5 Cool Things I Did With Local Language Models
Discover five unique, practical, and incredibly fun projects you can build on consumer hardware using open source local LLMs.
Top 5 Small Language Models for Agentic Tool Calling
Explore 5 small language models (SLMs) that feature open weights and first-class tool-calling support for local agentic workflows.