If your company handles proprietary code, financial data, or user records, sending any of it to a closed cloud API is a compliance question before it's a technical one. Until open reasoning weights showed up, that meant a real tradeoff: private models that could follow instructions but couldn't reason through anything hard, while genuine chain-of-thought reasoning stayed locked behind expensive hosted clusters.
That's no longer true. Distilled DeepSeek-R1 and quantized Qwen models let you run real multi-step reasoning, math, and code refactoring entirely offline, on hardware you already own.
Picking the Right Size for What You Actually Have
Reasoning models eat memory differently than a normal chat model, because they're generating an internal scratchpad before you ever see output, which fills the context window faster than you'd expect.
| Model | Quantization | Memory needed | Reasonable minimum hardware |
|---|---|---|---|
| DeepSeek-R1-Distill-Qwen-7B | Q4_K_M | Roughly 5-6 GB | An 8GB Apple Silicon Mac or an RTX 3060 |
| DeepSeek-R1-Distill-Qwen-14B | Q4_K_M | Roughly 10 GB | An RTX 4070 or a 16GB Apple Silicon Mac |
| DeepSeek-R1-Distill-Llama-70B | Q4_K_M | Roughly 42 GB | Dual RTX 3090s or an M-series Mac Studio |
| Full DeepSeek-R1 (671B MoE) | Q4_K_M | 400+ GB | Multi-GPU enterprise hardware, not a workstation |
For most individual engineers, the 14B distilled model is the sweet spot, enough reasoning depth to be genuinely useful, without needing hardware you don't already have.
Getting It Running
Pull the model and test it interactively first:
ollama --version
ollama pull deepseek-r1:14b
ollama run deepseek-r1:14b
Left on default settings, local reasoning models tend to ramble inside their <think> tags far longer than they need to. A custom Modelfile fixes that:
FROM deepseek-r1:14b
PARAMETER temperature 0.6
PARAMETER top_p 0.95
PARAMETER num_ctx 8192
SYSTEM """You are a senior systems engineer. Solve problems using
concise step-by-step logic. When writing code, provide working
implementations without conversational fluff."""
ollama create custom-reasoner -f Modelfile
Once it's built, you can talk to it through the same OpenAI-compatible client you'd use for a cloud model, just pointed at localhost:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama", # required by the client, ignored locally
)
response = client.chat.completions.create(
model="custom-reasoner",
messages=[{
"role": "user",
"content": "Write an idempotent SQL migration that adds a "
"nullable UUID column to a high-traffic Postgres "
"table with zero table locking.",
}],
)
print(response.choices[0].message.content)
The Three Ways This Actually Breaks
Push temperature above 0.8 and you'll watch the model get stuck oscillating between two candidate answers inside its thinking block until it burns through the entire context window. Keep it between 0.5 and 0.6, that range is where these models actually converge instead of loop.
When your prompt plus its reasoning chain exceeds num_ctx, Ollama either truncates the early context or spills to system swap, and generation speed can drop to something closer to a token a second. If you're seeing that, raise num_ctx before you blame the model.
And the smaller distilled variants, the 1.5B and 7B, occasionally slip into mixed-language scratchpad reasoning on harder problems, English and Mandarin blending mid-thought. Moving up to 14B mostly resolves it.
Frequently Asked Questions
Does this actually work with no internet connection at all?
Yes. Once the weights are downloaded once, you can disconnect entirely and inference keeps running, nothing about generation requires a live connection.
How does the 14B model actually compare to GPT-4o?
On algorithmic and logic-heavy tasks, it holds up well, sometimes matches or beats it. On broad world knowledge and creative writing, a large cloud model still has the edge, that gap doesn't close with quantization alone.
Is this hard on my GPU long-term?
No harder than any other sustained workload. As long as your cooling keeps temperatures under control, there's no meaningful extra wear from running local inference versus anything else that loads the GPU. If you need this wired into an internal tool with proper access control, not just a personal terminal session, that's a build worth doing right the first time. [SmartBuddy sets up local AI deployments for teams →](https://smartbuddy.cloud/start-a-project.html)
Comments
Comments are reviewed before appearing publicly.
No comments yet — be the first.