Hardware guide · DeepSeek-R1-Distill
What hardware do you need to run DeepSeek-R1-Distill-Qwen 32B?
DeepSeek-R1-Distill-Qwen 32B needs about 22.1 GB to run at Q4_K_M, so the lightest hardware that runs it is Nvidia GeForce RTX 4090 (24GB). 12 of 43 devices tested can run it.
- Needs (Q4_K_M)
- ~22.1 GB
- Devices that run it
- 12
- Too small
- 31
- Lightest that runs it
- ~23 GB
Runs at Q4_K_M using ~22.1 GB of ~31 GB usable.
What kind of hardware runs DeepSeek-R1-Distill-Qwen 32B
Of the 12 devices that run it, here is the split by hardware class.
Hardware that runs DeepSeek-R1-Distill-Qwen 32B
Ranked by usable memory, lightest first. Prices are approximate street prices for the device itself (a GPU is the card alone; a Mac is the whole machine). Tok/s is a bandwidth estimate; see methodology.
- TightNvidia GeForce RTX 4090 (24GB) from $1,599~23 GB usable · uses ~22.1 GB · ~33 tok/s est.
- TightNvidia GeForce RTX 3090 (24GB) from $1,499~23 GB usable · uses ~22.1 GB · ~31 tok/s est.
- TightAMD Radeon RX 7900 XTX (24GB) from $999~23 GB usable · uses ~22.1 GB · ~31 tok/s est.
- Yes32GB RAM Laptop (CPU/iGPU only) from $1,100~28 GB usable · uses ~22.1 GB · ~2 tok/s est.
- YesNvidia GeForce RTX 5090 (32GB) from $1,999~31 GB usable · uses ~22.1 GB · ~59 tok/s est.
- YesApple M4 Pro (48GB) from $2,399~32 GB usable · uses ~22.1 GB · ~11 tok/s est.
- YesApple M5 Pro (48GB) from $2,199~32 GB usable · uses ~22.1 GB · ~12 tok/s est.
- YesApple M4 Max (64GB) from $3,499~48 GB usable · uses ~22.1 GB · ~22 tok/s est.
- YesApple M4 Max (128GB) from $3,499~96 GB usable · uses ~22.1 GB · ~22 tok/s est.
- YesAMD Ryzen AI Halo (128GB) from $3,999~96 GB usable · uses ~22.1 GB · ~8 tok/s est.
- YesApple M5 Max (128GB) from $3,599~96 GB usable · uses ~22.1 GB · ~25 tok/s est.
- YesApple M3 Ultra (256GB) from $3,999~192 GB usable · uses ~22.1 GB · ~33 tok/s est.
Too small for DeepSeek-R1-Distill-Qwen 32B
See the full DeepSeek-R1-Distill-Qwen 32B memory breakdown, or compare it against other models.
Sources
-
github.com · 1 source
-
huggingface.co · 1 source
-
ollama.com · 1 source
Memory at Q4_K_M (22.1 GB = 19.85 GB weights + KV cache + overhead), Catalog updated 2026-10-05. See methodology.