Running Claude Code with a Local LLM: A Step-by-Step Guide
- 10 min read
Last updated: September 2026
You don’t need a separate server project, a Python virtual environment, or a proxy layer to run Claude Code against a local model. Ollama ships a native Anthropic-compatible API endpoint, so the whole setup is two environment variables.
Why Run Claude Code Against a Local LLM?
- Privacy – Code and prompts never leave your machine.
- No API costs – You’re paying in electricity and hardware, not per-token.
- Offline capability – Works with no internet connection.
- Model choice – Swap in whatever open-weight coding model fits your hardware.
The Setup
Install Ollama, pull a coding model, and point Claude Code at it:
ollama pull qwen3-coder:30b
export ANTHROPIC_BASE_URL=http://localhost:11434
export ANTHROPIC_AUTH_TOKEN=ollama
claude
That’s the entire configuration. No proxy, no separate API server, no config.yaml. qwen3-coder:30b is a reasonable default if you have 24GB+ of unified memory or VRAM; for lighter hardware, see the hardware benchmarks piece for which model actually fits your tier, and the full 2026 setup guide for backend options beyond Ollama.
What You Can Do With It
- Generate code: “Write a function that sorts a list using quicksort.”
- Debug locally: paste a function and ask why it’s returning the wrong value.
- Optimize queries: “Optimize this SQL query for performance.”
- Write and run tests: have it generate a unit test and fix what it finds.
- Work fully offline: no internet connection required once the model is pulled.
Where This Breaks Down
A local model isn’t a drop-in replacement for Claude’s API in every case. Quality-sensitive debugging, subtle logic errors, and anything requiring screenshot review are still where cloud models pull ahead. If you’re trying to decide whether local is worth it for your usage pattern rather than just how to set it up, see the local-vs-cloud decision framework.
Making infrastructure calls like this without a technical leader in the room is exactly how startups end up stuck with the wrong stack. If that sounds familiar, see when it makes sense to bring in a fractional CTO.
Want to explore how local LLMs can transform your development workflow? Let's discuss your AI integration needs!
Schedule Your Free Consultation