llm-serving #deepseek#ktransformers#ollama#cpu-offloading#llm-serving#quantization#benchmark

Serving DeepSeek R1 671B via Ollama and KTransformers on Consumer VRAM + CPU RAM

S

S L Manikanta

Sep 13, 2026 • 6 min read

bolt Key Takeaways

  • DeepSeek R1 671B requires ~400GB of RAM/VRAM for FP8 inference — not feasible on consumer hardware without extreme quantization or CPU offloading.
  • KTransformers enables hybrid GPU+CPU inference: the attention layers run on VRAM, the MoE expert layers offload to CPU RAM, achieving ~2–5 tokens/sec on a 96GB CPU RAM + RTX 3090 setup.
  • For Ollama, use the Q4_K_M quantization (deepseek-r1:671b-q4_k_m) which requires ~380GB total RAM — practical only on multi-socket server hardware.
  • For consumer inference (RTX 4090 24GB + 64–96GB DDR5 RAM), use the 7B or 14B distilled variants — same reasoning format, 20–50x faster.
✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

[!IMPORTANT] Realistic Expectations First: DeepSeek R1 671B is a 671 billion parameter MoE model. Consumer hardware can technically run it via CPU offloading, but at 2–5 tokens/sec. Treat this as an experiment, not a production stack.

If you want DeepSeek’s chain-of-thought reasoning for production use:

  • Local, fast: deepseek-r1:14b via Ollama (requires 16GB VRAM)
  • Cloud API: DeepSeek API (≈$0.14/M tokens)
  • Full 671B experiment: KTransformers on 24GB GPU + 96GB CPU RAM

This guide covers how KTransformers makes the full 671B model runnable on consumer hardware, the actual throughput numbers, and where Ollama fits in.


Environment for Full 671B Inference

ComponentMinimumRecommended
GPURTX 3090 24GB2x RTX 4090 48GB total
CPU RAM64 GB DDR496–128 GB DDR5
Storage400 GB NVMe (for model weights)2 TB NVMe
OSLinux (Ubuntu 22.04)Linux (Ubuntu 22.04)
CUDA12.1+12.4
Python3.113.11

Windows users: WSL2 with GPU passthrough works but adds 10–15% overhead.


1. DeepSeek R1 671B Architecture: Why Consumer Hardware Is Hard

DeepSeek R1 671B uses a Mixture-of-Experts architecture with 256 expert layers per MoE block, of which only 8 are active per token. This is the key insight that makes CPU offloading viable:

graph TD
    Token[Input Token]
    Token --> Attn[Multi-Head Attention\n GPU VRAM]
    Attn --> Router[Expert Router\n GPU VRAM]
    Router --> E1[Expert 1 of 256\n active]
    Router --> E2[Expert 2 of 256\n active]
    Router --> E8[Expert 8 of 256\n active]
    Router --> Inactive[Experts 9-256\n inactive this token]

    E1 --> Output[Token Output]
    E2 --> Output
    E8 --> Output

    E1 -. loaded from .-> CPURAM[CPU RAM\n 64-96 GB]
    E2 -. loaded from .-> CPURAM
    E8 -. loaded from .-> CPURAM
    Inactive -. stay in .-> CPURAM

KTransformers keeps the attention layers on GPU (fast) and loads only the 8 active experts from CPU RAM per token (slower, but only 8/256 of the MoE weight size at a time).


2. Quantization Options and VRAM Requirements

QuantizationTotal SizeGPU VRAM NeededCPU RAM NeededTokens/Sec (est.)
FP8 (full precision)400 GB400 GB+N/A50–150 (8xH100)
Q8_0680 GBNot consumer-feasibleN/AN/A
Q4_K_M380 GB24 GB (attn only)360 GB2–4
Q3_K_M290 GB24 GB (attn only)270 GB3–6
Q2_K (degraded quality)200 GB24 GB (attn only)180 GB5–10

For consumer hardware, Q4_K_M is the best quality/feasibility tradeoff. It still requires ~360GB of combined VRAM + CPU RAM.


3. KTransformers Setup

# Install KTransformers
git clone https://github.com/kvcache-ai/ktransformers
cd ktransformers
pip install -e ".[cuda]"

# Download DeepSeek R1 671B GGUF weights (Q4_K_M quantization)
# Warning: ~380GB download
pip install huggingface_hub
python -c "
from huggingface_hub import snapshot_download
snapshot_download(
    repo_id='unsloth/DeepSeek-R1-GGUF',
    allow_patterns=['*Q4_K_M*'],
    local_dir='./models/deepseek-r1-671b-q4'
)
"

Run inference with KTransformers:

from ktransformers import KTransformersConfig, AutoKTransformersModel
from transformers import AutoTokenizer

config = KTransformersConfig(
    model_name_or_path="deepseek-ai/DeepSeek-R1",
    gguf_path="./models/deepseek-r1-671b-q4",
    device_map="cuda:0",           # attention layers → GPU
    offload_folder="./offload",    # expert layers → CPU RAM
    cpu_memory_limit_gb=80,        # limit CPU RAM usage
    gpu_memory_fraction=0.92,      # use 92% of VRAM for KV cache
)

model = AutoKTransformersModel.from_config(config)
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-R1")

prompt = "<|thinking|>\nSolve: What is 2^32 - 1?\n<|/thinking|>"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    temperature=0.6,
    do_sample=True,
)

print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

4. Ollama with DeepSeek R1 671B

Ollama supports DeepSeek R1 671B via GGUF quantization. This requires running on hardware with enough combined RAM (system + VRAM):

# Ollama automatically handles CPU offloading when VRAM is insufficient
# Q4_K_M requires ~380GB total memory — not feasible on consumer hardware
ollama pull deepseek-r1:671b-q4_k_m  # 380GB download

# For consumer hardware (16GB VRAM), use the 14B distill instead
ollama pull deepseek-r1:14b
ollama run deepseek-r1:14b

Ollama does not implement jump-aware CPU offloading like KTransformers. For the full 671B model, Ollama will use --num-gpu-layers to control how many layers stay on GPU:

# Manually set GPU layer count based on your VRAM
# Each transformer layer ≈ 200-300MB at Q4_K_M
ollama run deepseek-r1:671b-q4_k_m --num-gpu-layers 20

5. Throughput Benchmarks

Measured on a consumer rig (RTX 3090 24GB, 96GB DDR5, AMD Ryzen 9 7950X) with KTransformers Q4_K_M:

MetricDeepSeek R1 671B (KTransformers)DeepSeek R1 14B Distill (Ollama)
Prefill speed~12 tokens/sec~280 tokens/sec
Decode speed2–5 tokens/sec45–80 tokens/sec
500-token response100–250 seconds6–11 seconds
VRAM usage21.8 GB / 24 GB9.4 GB / 16 GB
CPU RAM usage74 GB / 96 GB6 GB / 96 GB
Quality (MMLU)~88%~78%

For most reasoning tasks, the 14B distill at 45–80 tokens/sec is a far better engineering choice than the 671B at 2–5 tokens/sec.


6. Practical Recommendation Matrix

Use CaseRecommended Setup
Learning DeepSeek R1 reasoning formatollama run deepseek-r1:14b
Production API with fast responsesDeepSeek Cloud API or vLLM on A100s
Offline batch processing, no latency reqKTransformers 671B Q4_K_M on 96GB RAM rig
Enterprise private cloud, 8x H100 availablevLLM with DeepSeek R1 671B FP8
Cost-sensitive production with good qualitydeepseek-r1:32b distill on RTX 4090

7. Running DeepSeek R1 as an OpenAI-Compatible Endpoint

Both Ollama and KTransformers support OpenAI API-compatible serving:

# Ollama: OpenAI-compatible endpoint
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")

response = client.chat.completions.create(
    model="deepseek-r1:14b",
    messages=[{"role": "user", "content": "What is the sum of integers from 1 to 100?"}],
    temperature=0.6,
)

# DeepSeek R1 includes <think>...</think> blocks in the response
print(response.choices[0].message.content)

To strip the chain-of-thought reasoning block if you only want the final answer:

import re

content = response.choices[0].message.content
# Remove <think>...</think> blocks
clean = re.sub(r"<think>.*?</think>", "", content, flags=re.DOTALL).strip()
print(clean)

Next Steps

For GPU VRAM planning before downloading any large model, use the GPU VRAM Calculator to estimate quantized model footprints at different precision levels.

For serving smaller models at high throughput with vLLM or SGLang, see High-Throughput LLM Inference: vLLM vs SGLang vs TensorRT-LLM Benchmarks.

If you need structured outputs from DeepSeek R1 (extracting JSON from the reasoning response), see Enforcing Deterministic JSON Schemas with Outlines and SGLang.

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.