Run Claude Code Locally with Ollama (2026 Setup)
In this article
- Why Run an AI Coding Agent Locally?
- Prerequisites
- Step 1: Install Ollama
- Step 2: Pull and Test the Model
- Step 3: Set an Appropriate Context Length
- Step 4: Install Claude Code
- Step 5: Connect Claude Code to Ollama
- Step 6: Run Your First Agentic Task
- Bonus: Use a Local GGUF File Directly
- Which Models Work Best?
- What You've Built
- What to Read Next
Why Run an AI Coding Agent Locally?
Most terminal-based coding agents phone home to a cloud API. That works fine until you're on a flaky connection, need to keep proprietary code off external servers, or simply want to avoid per-token billing on long agentic sessions.
Running Claude Code against a local Ollama server can keep model prompts and inference on your machine while preserving Claude Code's terminal workflow. Ollama exposes an Anthropic-compatible Messages API, which is the compatibility layer Claude Code uses here; this is not an OpenAI /v1 endpoint. This guide walks through installation, a 64K context configuration, routing verification, and a multi-step coding test.
Prerequisites
Before you start, make sure your system has:
Hardware:
- Enough GPU or system memory for the Ollama model you select
- Disk space for the model weights plus your project
- A 64K context configuration for reliable agentic sessions
If you have no GPU, inference will fall back to CPU, which is noticeably slower but still functional.
Software:
- Linux, macOS, or Windows; WSL2 is optional if it already matches your development workflow
- A working GPU driver if you plan to use GPU acceleration
- Node.js (required for the Claude Code installer)
Verify your GPU setup before proceeding:
nvidia-smi
You should see your GPU model, available VRAM, and the active CUDA version listed in the output.
Step 1: Install Ollama
Ollama is the local runtime that downloads, manages, and serves models. It also exposes an HTTP API that external tools — including Claude Code — can talk to.
On Linux, a single command handles the full installation:
curl -fsSL https://ollama.com/install.sh | sh
On macOS and Windows, download the installer from ollama.com and follow the on-screen instructions. Ollama runs as a background service and checks for updates automatically.
Once installed, confirm the version:
ollama -v
If you get a "command not found" error, the service may not have started yet. Launch it manually in one terminal:
ollama serve
Then run ollama -v in a second terminal. Once the version prints cleanly, Ollama is ready.
Step 2: Pull and Test the Model
With Ollama running, download a local model. Ollama's current Claude Code integration guide recommends glm-4.7-flash and qwen3.5 for local use. This walkthrough uses GLM 4.7 Flash; choose Qwen3.5 if it is a better fit for your hardware and task mix.
ollama pull glm-4.7-flash
After the download completes, run a quick sanity check in interactive mode:
ollama run glm-4.7-flash
Type a short prompt and verify you get a sensible response without an API or model-loading error. Record the first-token latency on your machine; it varies with model size, quantization, available memory, and whether the weights were already loaded.
You can also verify the model responds over the local HTTP API, which is how Claude Code will communicate with it:
curl http://localhost:11434/api/chat -d '{
"model": "glm-4.7-flash",
"messages": [{"role": "user", "content": "Hello!"}]
}'
A JSON response confirms the API is live and the model is loaded.
Step 3: Set an Appropriate Context Length
Claude Code and agentic workflows in general require a reasonable context window to function well. However, very large windows can cause two problems: inference slows dramatically, and the model sometimes enters repetitive thinking loops.
Ollama recommends at least 64K tokens for coding agents. A smaller window can truncate tool results, project context, or the conversation state during a multi-file task. The trade-off is memory: if 64K does not fit, use a smaller model rather than silently shrinking the context and assuming equivalent behavior.
Stop the running Ollama server with Ctrl + C, then restart it with the context length override:
OLLAMA_CONTEXT_LENGTH=65536 ollama serve
Confirm the setting is active in a new terminal:
ollama ps
The CONTEXT column should read 65536 for the loaded model. The model ID, size, and processor columns vary by Ollama build and hardware, so use the context value—not a copied sample row—as the acceptance check.
Step 4: Install Claude Code
Claude Code is a terminal-based coding agent built by Anthropic. You interact with it in natural language, and it handles writing, editing, refactoring, and executing code as part of multi-step workflows.
Install it with the official script:
curl -fsSL https://claude.ai/install.sh | bash
Once complete, the claude command will be available in your terminal.
Step 5: Connect Claude Code to Ollama
Navigate to your project directory:
mkdir my-local-project
cd my-local-project
The recommended way to launch Claude Code with Ollama is to use the built-in launch command, which automatically configures the API routing for you:
ollama launch claude --model glm-4.7-flash
Alternatively, configure the environment variables manually and run claude directly. ANTHROPIC_BASE_URL ends at port 11434; adding /v1 routes to the wrong compatibility surface.
# Linux / macOS
export ANTHROPIC_BASE_URL="http://localhost:11434"
export ANTHROPIC_AUTH_TOKEN="ollama"
export ANTHROPIC_API_KEY=""
claude
# Windows (PowerShell)
$env:ANTHROPIC_BASE_URL = "http://localhost:11434"
$env:ANTHROPIC_AUTH_TOKEN = "ollama"
$env:ANTHROPIC_API_KEY = ""
claude
Once the Claude Code interface opens in your terminal, confirm it is pointing to your local model:
/model
If the output shows glm-4.7-flash, confirm the running context in a second terminal:
ollama ps
The setup is complete only when both the selected model and the 65536-token context are correct. If requests still reach Anthropic, start a clean shell, re-export the variables, and launch Claude Code again from that shell.
Step 6: Run Your First Agentic Task
With everything wired up, Claude Code will respond using your locally running model. Start with a simple greeting to confirm routing and response time before allowing file writes or shell commands.
For a more realistic test, ask Claude Code to build something complete:
"Build a command-line Snake game in Python."
Before it generates code, enable Planning Mode by pressing Shift + Tab twice. The model will outline its approach first — you can review the plan, ask for adjustments, and then tell it to proceed. Claude Code will create the required files and provide instructions for running the result.
Bonus: Use a Local GGUF File Directly
If you already have a GGUF model file downloaded and want to skip re-downloading via Ollama, you can register it manually with a Modelfile:
FROM ./glm-4.7-flash.gguf
PARAMETER temperature 0.8
PARAMETER top_p 0.95
PARAMETER repeat_penalty 1.0
Register it:
ollama create glm-4.7-flash-local -f Modelfile
Then run it like any other Ollama model:
ollama run glm-4.7-flash-local
Which Models Work Best?
Not every local model handles agentic workflows cleanly. Tool calling, long conversations, and multi-step planning put real demands on instruction following. These are the current choices listed by Ollama for Claude Code:
| Model | Execution | Use it when |
|---|---|---|
glm-4.7-flash | Local | You want the documented local starting point and the weights fit your machine. |
qwen3.5 | Local | You want a second local option for comparing tool use and code quality. |
kimi-k2.5:cloud, glm-5:cloud, minimax-m2.7:cloud, qwen3.5:cloud | Ollama cloud | You need a larger hosted model and accept that inference is no longer local. |
Start with one local model, clone a disposable repository, and run the same edit-and-test task several times. A model is suitable only if it consistently chooses the right files, emits valid tool calls, and recovers from command failures—not merely because it can answer a coding question.
What You've Built
At the end of this setup you have Claude Code routed to a local model endpoint that:
- Runs model inference on your own hardware when a non-cloud Ollama model is selected
- Uses the same Claude Code interface and workflow patterns as the cloud version
- Can be pointed at any Ollama-compatible model with a single flag change
For sensitive codebases, verify the selected model is local, inspect the process with ollama ps, and apply your normal outbound-network controls. That is a stronger privacy check than relying on a launch command alone.
What to Read Next
Keep building with local AI:
- How to Run Local LLMs Securely Using Ollama — Secure your Ollama instance, network-bind it correctly, and avoid exposing models unintentionally
- Local RAG Tutorial: LangChain, Ollama & ChromaDB with Ragas — Combine Ollama with a vector store to build a fully offline retrieval-augmented pipeline
- Building AI Agents with Local Small Language Models — Move beyond single-turn completions — tool calling and multi-step agent loops on your own hardware
- Run Claude Code for Free with OpenRouter — No GPU? Route Claude Code through high-limit free-tier models on OpenRouter instead
Frequently asked questions
Can you run Claude Code with a local model instead of Anthropic's API?
Yes. Ollama implements the Anthropic Messages API that Claude Code expects. Use `ollama launch claude` or set `ANTHROPIC_BASE_URL` to `http://localhost:11434`; do not append `/v1` for this integration.
Which Ollama models work best with Claude Code for coding tasks?
Ollama currently recommends `glm-4.7-flash` and `qwen3.5` for local use, with Kimi K2.5, GLM-5, MiniMax M2.7, and Qwen3.5 cloud variants also supported. Start with a model that fits your hardware and verify file edits, shell calls, and long-context behavior on a disposable repository.
Do I need a GPU to run Claude Code with Ollama?
No, but CPU inference makes multi-step agent sessions substantially slower. GPU memory requirements depend on the exact Ollama model and quantization, so check the model size before downloading rather than relying on a single VRAM estimate.
Is running Claude Code locally actually private?
Model inference and prompts stay local when Claude Code points to a local Ollama server and you select a local model. Claude Code is still a separate Anthropic application, and cloud-tagged Ollama models use remote infrastructure, so audit the selected model and your network policy before making an absolute privacy claim.
Related Guides
Building AI Agents with Local Small Language Models (SLMs)
Learn how to build fully functional, private AI agents on your own hardware using Ollama and LangChain with lightweight models under 10B parameters.

Structured Output with Local LLMs: When Valid JSON Is Not Enough
Gemma 4's 4B model returns schema-valid JSON that still includes the wrong device. Here's the decomposition pattern that fixes it.
The Complete Developer Guide to Running LLMs Locally: From Ollama to Production
Everything you need to run LLMs on your own hardware in 2026: VRAM sizing, model formats, an 8-tool comparison table, a full local RAG pipeline, and Docker production deployment with GPU passthrough and Nginx auth.