Can I run this AI model? GPU memory (VRAM) calculator
Pick a model and a precision: see how much GPU memory it needs, which ComputeBD AI Computer fits, how fast it runs, and the price per hour in taka.
You need about 17.7 GB of GPU memory to run Llama 3.1 8B at FP16 / BF16.
- Model weights 15 GB
- Conversation memory (KV cache) 0.5 GB
- Runtime & working space 2.2 GB
NVIDIA L4 is the most affordable AI Computer that fits — ৳147/hour, billed by the minute.
| AI Computer | Fits? | Speed (1 user) | Price |
|---|---|---|---|
| NVIDIA L424 GB | Fits |
≈ 11 tokens/s | ৳147/hr |
| NVIDIA A1024 GB | Fits |
≈ 22 tokens/s | ৳196/hr |
| NVIDIA L40S48 GB | Fits |
≈ 32 tokens/s | ৳343/hr |
| NVIDIA A10080 GB | Fits |
≈ 76 tokens/s | ৳638/hr |
| NVIDIA H10080 GB | Fits |
≈ 125 tokens/s | ৳1,047/hr |
| NVIDIA H200141 GB | Fits |
≈ 179 tokens/s | ৳1,390/hr |
| NVIDIA B200180 GB | Fits |
≈ 299 tokens/s | ৳1,652/hr |
| NVIDIA B300288 GB | Fits |
≈ 299 tokens/s | ৳1,880/hr |
Quick answers: GPU memory for popular models
| Model | FP16 | 8-bit | 4-bit | LoRA fine-tune |
|---|---|---|---|---|
| Llama 3.1 8B | 17.7 GB | 10.1 GB | 6 GB | 18.9 GB |
| Llama 3.2 3B | 7.9 GB | 4.9 GB | 3.3 GB | 8.4 GB |
| Llama 3.1 70B | 144.3 GB | 77.7 GB | 42.2 GB | 154.1 GB |
| Qwen2.5 7B | 16.5 GB | 9.4 GB | 5.5 GB | 18 GB |
| Qwen2.5 14B | 31.5 GB | 17.6 GB | 10.1 GB | 33.5 GB |
| Qwen2.5 32B | 68 GB | 37.1 GB | 20.6 GB | 72.4 GB |
| Qwen2.5 72B | 148.5 GB | 79.9 GB | 43.4 GB | 158.7 GB |
| Mistral 7B | 16.1 GB | 9.2 GB | 5.6 GB | 17.2 GB |
| Mixtral 8x7B | 95.4 GB | 51.4 GB | 27.9 GB | 102.4 GB |
| Gemma 2 9B | 20.9 GB | 12.2 GB | 7.5 GB | 21.5 GB |
| Gemma 2 27B | 57.2 GB | 31.5 GB | 17.8 GB | 60.3 GB |
| Phi-3.5 mini (3.8B) | 10.2 GB | 6.6 GB | 4.7 GB | 9.8 GB |
Inference with a 4k context for one user. Click a model for its own page.
How the estimate works
- Weights = parameters × bytes per parameter (FP16 = 2 bytes, 8-bit ≈ 1, 4-bit ≈ 0.56).
- KV cache (conversation memory) grows with context length and the number of people at once: 2 × layers × KV-heads × head size × tokens × 2 bytes.
- Runtime: about 1 GB for CUDA plus ~8 % of the weights for working space.
- Fine-tuning: LoRA keeps an FP16 copy of the model and trains ~1 % extra parameters; QLoRA keeps a 4-bit copy; full fine-tuning needs ~16 bytes per parameter. Gradient checkpointing is assumed.
- Speed: generating text for one user is limited by memory bandwidth, so tokens/second ≈ bandwidth × 60 % ÷ model size.
These are good planning estimates, not guarantees: the real figure depends on the software (vLLM, llama.cpp, Transformers…), settings and your data. Leave ~10 % spare.
VRAM needed, model by model
Frequently asked questions
How much VRAM does Llama 3.1 8B need?
About 17.7 GB at FP16 and 6 GB at 4-bit — a 24 GB L4 runs it at full quality.
Does 4-bit make the model worse?
Slightly. For chat and most tasks the difference is small, and it needs about a quarter of the memory. For precise work (code, maths, evaluation) prefer 8-bit or FP16.
Why does more context need more memory?
The model keeps a KV cache for every token it is considering. Long documents or many users at once multiply it.
Try it on a real AI Computer
JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.