How much VRAM does Llama 3.2 3B need?
Llama 3.2 3B needs about 7.9 GB of GPU memory at FP16 and about 3.3 GB at 4-bit. Pick a precision below to see which ComputeBD AI Computer fits and what it costs per hour in taka.
You need about 7.9 GB of GPU memory to run Llama 3.2 3B at FP16 / BF16.
- Model weights 6 GB
- Conversation memory (KV cache) 0.4 GB
- Runtime & working space 1.5 GB
NVIDIA L4 is the most affordable AI Computer that fits — ৳147/hour, billed by the minute.
| AI Computer | Fits? | Speed (1 user) | Price |
|---|---|---|---|
| NVIDIA L424 GB | Fits |
≈ 28 tokens/s | ৳147/hr |
| NVIDIA A1024 GB | Fits |
≈ 56 tokens/s | ৳196/hr |
| NVIDIA L40S48 GB | Fits |
≈ 81 tokens/s | ৳343/hr |
| NVIDIA A10080 GB | Fits |
≈ 191 tokens/s | ৳638/hr |
| NVIDIA H10080 GB | Fits |
≈ 313 tokens/s | ৳1,047/hr |
| NVIDIA H200141 GB | Fits |
≈ 449 tokens/s | ৳1,390/hr |
| NVIDIA B200180 GB | Fits |
≈ 748 tokens/s | ৳1,652/hr |
| NVIDIA B300288 GB | Fits |
≈ 748 tokens/s | ৳1,880/hr |
Llama 3.2 3B at each precision
| Precision | GPU memory needed | Cheapest ComputeBD AI Computer |
|---|---|---|
| FP32full precision — rarely needed | 14.4 GB | NVIDIA L4 · ৳147/hr |
| FP16 / BF16standard — original quality | 7.9 GB | NVIDIA L4 · ৳147/hr |
| 8-bitalmost no quality loss | 4.9 GB | NVIDIA L4 · ৳147/hr |
| 4-bit (GGUF / AWQ / GPTQ)small quality loss, ~4× less memory | 3.3 GB | NVIDIA L4 · ৳147/hr |
- Made by
- Meta
- Parameters
- 3.21B
- Max context
- 131,072 tokens
- LoRA / QLoRA fine-tune
- ≈ 8.4 GB / 4.1 GB
- Hugging Face
- meta-llama/Llama-3.2-3B-Instruct
How the estimate works
- Weights = parameters × bytes per parameter (FP16 = 2 bytes, 8-bit ≈ 1, 4-bit ≈ 0.56).
- KV cache (conversation memory) grows with context length and the number of people at once: 2 × layers × KV-heads × head size × tokens × 2 bytes.
- Runtime: about 1 GB for CUDA plus ~8 % of the weights for working space.
- Fine-tuning: LoRA keeps an FP16 copy of the model and trains ~1 % extra parameters; QLoRA keeps a 4-bit copy; full fine-tuning needs ~16 bytes per parameter. Gradient checkpointing is assumed.
- Speed: generating text for one user is limited by memory bandwidth, so tokens/second ≈ bandwidth × 60 % ÷ model size.
These are good planning estimates, not guarantees: the real figure depends on the software (vLLM, llama.cpp, Transformers…), settings and your data. Leave ~10 % spare.
VRAM needed, model by model
Frequently asked questions
How much VRAM does Llama 3.2 3B need?
About 7.9 GB at FP16 (original quality) and about 3.3 GB at 4-bit, for one user with a 4k context.
Which GPU can run Llama 3.2 3B?
At full quality it fits on an NVIDIA L4 (24 GB), ৳147/hour on ComputeBD.
How much VRAM to fine-tune Llama 3.2 3B?
About 8.4 GB with LoRA and 4.1 GB with QLoRA (sequence length 2,048, batch 1, gradient checkpointing). Estimate the time and cost with the fine-tuning cost estimator.
Try it on a real AI Computer
JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.