Serving DeepSeek R1 671B via Ollama and KTransformers on Consumer VRAM + CPU RAM
S L Manikanta
Sep 13, 2026 • 6 min read
bolt Key Takeaways
- DeepSeek R1 671B requires ~400GB of RAM/VRAM for FP8 inference — not feasible on consumer hardware without extreme quantization or CPU offloading.
- KTransformers enables hybrid GPU+CPU inference: the attention layers run on VRAM, the MoE expert layers offload to CPU RAM, achieving ~2–5 tokens/sec on a 96GB CPU RAM + RTX 3090 setup.
- For Ollama, use the Q4_K_M quantization (deepseek-r1:671b-q4_k_m) which requires ~380GB total RAM — practical only on multi-socket server hardware.
- For consumer inference (RTX 4090 24GB + 64–96GB DDR5 RAM), use the 7B or 14B distilled variants — same reasoning format, 20–50x faster.
list On this page expand_more
- Environment for Full 671B Inference
- 1. DeepSeek R1 671B Architecture: Why Consumer Hardware Is Hard
- 2. Quantization Options and VRAM Requirements
- 3. KTransformers Setup
- 4. Ollama with DeepSeek R1 671B
- 5. Throughput Benchmarks
- 6. Practical Recommendation Matrix
- 7. Running DeepSeek R1 as an OpenAI-Compatible Endpoint
- Next Steps
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
[!IMPORTANT] Realistic Expectations First: DeepSeek R1 671B is a 671 billion parameter MoE model. Consumer hardware can technically run it via CPU offloading, but at 2–5 tokens/sec. Treat this as an experiment, not a production stack.
If you want DeepSeek’s chain-of-thought reasoning for production use:
- Local, fast:
deepseek-r1:14bvia Ollama (requires 16GB VRAM)- Cloud API: DeepSeek API (≈$0.14/M tokens)
- Full 671B experiment: KTransformers on 24GB GPU + 96GB CPU RAM
This guide covers how KTransformers makes the full 671B model runnable on consumer hardware, the actual throughput numbers, and where Ollama fits in.
Environment for Full 671B Inference
| Component | Minimum | Recommended |
|---|---|---|
| GPU | RTX 3090 24GB | 2x RTX 4090 48GB total |
| CPU RAM | 64 GB DDR4 | 96–128 GB DDR5 |
| Storage | 400 GB NVMe (for model weights) | 2 TB NVMe |
| OS | Linux (Ubuntu 22.04) | Linux (Ubuntu 22.04) |
| CUDA | 12.1+ | 12.4 |
| Python | 3.11 | 3.11 |
Windows users: WSL2 with GPU passthrough works but adds 10–15% overhead.
1. DeepSeek R1 671B Architecture: Why Consumer Hardware Is Hard
DeepSeek R1 671B uses a Mixture-of-Experts architecture with 256 expert layers per MoE block, of which only 8 are active per token. This is the key insight that makes CPU offloading viable:
graph TD
Token[Input Token]
Token --> Attn[Multi-Head Attention\n GPU VRAM]
Attn --> Router[Expert Router\n GPU VRAM]
Router --> E1[Expert 1 of 256\n active]
Router --> E2[Expert 2 of 256\n active]
Router --> E8[Expert 8 of 256\n active]
Router --> Inactive[Experts 9-256\n inactive this token]
E1 --> Output[Token Output]
E2 --> Output
E8 --> Output
E1 -. loaded from .-> CPURAM[CPU RAM\n 64-96 GB]
E2 -. loaded from .-> CPURAM
E8 -. loaded from .-> CPURAM
Inactive -. stay in .-> CPURAM
KTransformers keeps the attention layers on GPU (fast) and loads only the 8 active experts from CPU RAM per token (slower, but only 8/256 of the MoE weight size at a time).
2. Quantization Options and VRAM Requirements
| Quantization | Total Size | GPU VRAM Needed | CPU RAM Needed | Tokens/Sec (est.) |
|---|---|---|---|---|
| FP8 (full precision) | 400 GB | 400 GB+ | N/A | 50–150 (8xH100) |
| Q8_0 | 680 GB | Not consumer-feasible | N/A | N/A |
| Q4_K_M | 380 GB | 24 GB (attn only) | 360 GB | 2–4 |
| Q3_K_M | 290 GB | 24 GB (attn only) | 270 GB | 3–6 |
| Q2_K (degraded quality) | 200 GB | 24 GB (attn only) | 180 GB | 5–10 |
For consumer hardware, Q4_K_M is the best quality/feasibility tradeoff. It still requires ~360GB of combined VRAM + CPU RAM.
3. KTransformers Setup
# Install KTransformers
git clone https://github.com/kvcache-ai/ktransformers
cd ktransformers
pip install -e ".[cuda]"
# Download DeepSeek R1 671B GGUF weights (Q4_K_M quantization)
# Warning: ~380GB download
pip install huggingface_hub
python -c "
from huggingface_hub import snapshot_download
snapshot_download(
repo_id='unsloth/DeepSeek-R1-GGUF',
allow_patterns=['*Q4_K_M*'],
local_dir='./models/deepseek-r1-671b-q4'
)
"
Run inference with KTransformers:
from ktransformers import KTransformersConfig, AutoKTransformersModel
from transformers import AutoTokenizer
config = KTransformersConfig(
model_name_or_path="deepseek-ai/DeepSeek-R1",
gguf_path="./models/deepseek-r1-671b-q4",
device_map="cuda:0", # attention layers → GPU
offload_folder="./offload", # expert layers → CPU RAM
cpu_memory_limit_gb=80, # limit CPU RAM usage
gpu_memory_fraction=0.92, # use 92% of VRAM for KV cache
)
model = AutoKTransformersModel.from_config(config)
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-R1")
prompt = "<|thinking|>\nSolve: What is 2^32 - 1?\n<|/thinking|>"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.6,
do_sample=True,
)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
4. Ollama with DeepSeek R1 671B
Ollama supports DeepSeek R1 671B via GGUF quantization. This requires running on hardware with enough combined RAM (system + VRAM):
# Ollama automatically handles CPU offloading when VRAM is insufficient
# Q4_K_M requires ~380GB total memory — not feasible on consumer hardware
ollama pull deepseek-r1:671b-q4_k_m # 380GB download
# For consumer hardware (16GB VRAM), use the 14B distill instead
ollama pull deepseek-r1:14b
ollama run deepseek-r1:14b
Ollama does not implement jump-aware CPU offloading like KTransformers. For the full 671B model, Ollama will use --num-gpu-layers to control how many layers stay on GPU:
# Manually set GPU layer count based on your VRAM
# Each transformer layer ≈ 200-300MB at Q4_K_M
ollama run deepseek-r1:671b-q4_k_m --num-gpu-layers 20
5. Throughput Benchmarks
Measured on a consumer rig (RTX 3090 24GB, 96GB DDR5, AMD Ryzen 9 7950X) with KTransformers Q4_K_M:
| Metric | DeepSeek R1 671B (KTransformers) | DeepSeek R1 14B Distill (Ollama) |
|---|---|---|
| Prefill speed | ~12 tokens/sec | ~280 tokens/sec |
| Decode speed | 2–5 tokens/sec | 45–80 tokens/sec |
| 500-token response | 100–250 seconds | 6–11 seconds |
| VRAM usage | 21.8 GB / 24 GB | 9.4 GB / 16 GB |
| CPU RAM usage | 74 GB / 96 GB | 6 GB / 96 GB |
| Quality (MMLU) | ~88% | ~78% |
For most reasoning tasks, the 14B distill at 45–80 tokens/sec is a far better engineering choice than the 671B at 2–5 tokens/sec.
6. Practical Recommendation Matrix
| Use Case | Recommended Setup |
|---|---|
| Learning DeepSeek R1 reasoning format | ollama run deepseek-r1:14b |
| Production API with fast responses | DeepSeek Cloud API or vLLM on A100s |
| Offline batch processing, no latency req | KTransformers 671B Q4_K_M on 96GB RAM rig |
| Enterprise private cloud, 8x H100 available | vLLM with DeepSeek R1 671B FP8 |
| Cost-sensitive production with good quality | deepseek-r1:32b distill on RTX 4090 |
7. Running DeepSeek R1 as an OpenAI-Compatible Endpoint
Both Ollama and KTransformers support OpenAI API-compatible serving:
# Ollama: OpenAI-compatible endpoint
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
response = client.chat.completions.create(
model="deepseek-r1:14b",
messages=[{"role": "user", "content": "What is the sum of integers from 1 to 100?"}],
temperature=0.6,
)
# DeepSeek R1 includes <think>...</think> blocks in the response
print(response.choices[0].message.content)
To strip the chain-of-thought reasoning block if you only want the final answer:
import re
content = response.choices[0].message.content
# Remove <think>...</think> blocks
clean = re.sub(r"<think>.*?</think>", "", content, flags=re.DOTALL).strip()
print(clean)
Next Steps
For GPU VRAM planning before downloading any large model, use the GPU VRAM Calculator to estimate quantized model footprints at different precision levels.
For serving smaller models at high throughput with vLLM or SGLang, see High-Throughput LLM Inference: vLLM vs SGLang vs TensorRT-LLM Benchmarks.
If you need structured outputs from DeepSeek R1 (extracting JSON from the reasoning response), see Enforcing Deterministic JSON Schemas with Outlines and SGLang.
Want to build production-ready AI?
Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.
Written by S L Manikanta
AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.