← All free tools

How much VRAM does DeepSeek-R1-Distill-Qwen 14B need?

DeepSeek-R1-Distill-Qwen 14B needs about 31.5 GB of GPU memory at FP16 and about 10.1 GB at 4-bit. Pick a precision below to see which ComputeBD AI Computer fits and what it costs per hour in taka.

What do you want to do?

You need about 31.5 GB of GPU memory to run DeepSeek-R1-Distill-Qwen 14B at FP16 / BF16.

  • Model weights 27.6 GB
  • Conversation memory (KV cache) 0.8 GB
  • Runtime & working space 3.2 GB

NVIDIA L40S is the most affordable AI Computer that fits — ৳343/hour, billed by the minute.

AI ComputerFits?Speed (1 user)Price
NVIDIA L424 GB Too small
≈ 6 tokens/s ৳147/hr
NVIDIA A1024 GB Too small
≈ 12 tokens/s ৳196/hr
NVIDIA L40S48 GB Fits
≈ 18 tokens/s ৳343/hr
NVIDIA A10080 GB Fits
≈ 41 tokens/s ৳638/hr
NVIDIA H10080 GB Fits
≈ 68 tokens/s ৳1,047/hr
NVIDIA H200141 GB Fits
≈ 97 tokens/s ৳1,390/hr
NVIDIA B200180 GB Fits
≈ 162 tokens/s ৳1,652/hr
NVIDIA B300288 GB Fits
≈ 162 tokens/s ৳1,880/hr

DeepSeek-R1-Distill-Qwen 14B at each precision

PrecisionGPU memory neededCheapest ComputeBD AI Computer
FP32full precision — rarely needed61.3 GBNVIDIA A100 · ৳638/hr
FP16 / BF16standard — original quality31.5 GBNVIDIA L40S · ৳343/hr
8-bitalmost no quality loss17.6 GBNVIDIA L4 · ৳147/hr
4-bit (GGUF / AWQ / GPTQ)small quality loss, ~4× less memory10.1 GBNVIDIA L4 · ৳147/hr
Made by
DeepSeek
Parameters
14.8B
Max context
131,072 tokens
LoRA / QLoRA fine-tune
≈ 33.5 GB / 13.7 GB
Hugging Face
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B

How the estimate works

These are good planning estimates, not guarantees: the real figure depends on the software (vLLM, llama.cpp, Transformers…), settings and your data. Leave ~10 % spare.

VRAM needed, model by model

Frequently asked questions

How much VRAM does DeepSeek-R1-Distill-Qwen 14B need?

About 31.5 GB at FP16 (original quality) and about 10.1 GB at 4-bit, for one user with a 4k context.

Which GPU can run DeepSeek-R1-Distill-Qwen 14B?

At full quality it fits on an NVIDIA L40S (48 GB), ৳343/hour on ComputeBD.

How much VRAM to fine-tune DeepSeek-R1-Distill-Qwen 14B?

About 33.5 GB with LoRA and 13.7 GB with QLoRA (sequence length 2,048, batch 1, gradient checkpointing). Estimate the time and cost with the fine-tuning cost estimator.

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.