← All free tools

Can I run this AI model? GPU memory (VRAM) calculator

Pick a model and a precision: see how much GPU memory it needs, which ComputeBD AI Computer fits, how fast it runs, and the price per hour in taka.

What do you want to do?

You need about 17.7 GB of GPU memory to run Llama 3.1 8B at FP16 / BF16.

  • Model weights 15 GB
  • Conversation memory (KV cache) 0.5 GB
  • Runtime & working space 2.2 GB

NVIDIA L4 is the most affordable AI Computer that fits — ৳147/hour, billed by the minute.

AI ComputerFits?Speed (1 user)Price
NVIDIA L424 GB Fits
≈ 11 tokens/s ৳147/hr
NVIDIA A1024 GB Fits
≈ 22 tokens/s ৳196/hr
NVIDIA L40S48 GB Fits
≈ 32 tokens/s ৳343/hr
NVIDIA A10080 GB Fits
≈ 76 tokens/s ৳638/hr
NVIDIA H10080 GB Fits
≈ 125 tokens/s ৳1,047/hr
NVIDIA H200141 GB Fits
≈ 179 tokens/s ৳1,390/hr
NVIDIA B200180 GB Fits
≈ 299 tokens/s ৳1,652/hr
NVIDIA B300288 GB Fits
≈ 299 tokens/s ৳1,880/hr

Quick answers: GPU memory for popular models

ModelFP168-bit4-bitLoRA fine-tune
Llama 3.1 8B17.7 GB10.1 GB6 GB18.9 GB
Llama 3.2 3B7.9 GB4.9 GB3.3 GB8.4 GB
Llama 3.1 70B144.3 GB77.7 GB42.2 GB154.1 GB
Qwen2.5 7B16.5 GB9.4 GB5.5 GB18 GB
Qwen2.5 14B31.5 GB17.6 GB10.1 GB33.5 GB
Qwen2.5 32B68 GB37.1 GB20.6 GB72.4 GB
Qwen2.5 72B148.5 GB79.9 GB43.4 GB158.7 GB
Mistral 7B16.1 GB9.2 GB5.6 GB17.2 GB
Mixtral 8x7B95.4 GB51.4 GB27.9 GB102.4 GB
Gemma 2 9B20.9 GB12.2 GB7.5 GB21.5 GB
Gemma 2 27B57.2 GB31.5 GB17.8 GB60.3 GB
Phi-3.5 mini (3.8B)10.2 GB6.6 GB4.7 GB9.8 GB

Inference with a 4k context for one user. Click a model for its own page.

How the estimate works

These are good planning estimates, not guarantees: the real figure depends on the software (vLLM, llama.cpp, Transformers…), settings and your data. Leave ~10 % spare.

VRAM needed, model by model

Frequently asked questions

How much VRAM does Llama 3.1 8B need?

About 17.7 GB at FP16 and 6 GB at 4-bit — a 24 GB L4 runs it at full quality.

Does 4-bit make the model worse?

Slightly. For chat and most tasks the difference is small, and it needs about a quarter of the memory. For precise work (code, maths, evaluation) prefer 8-bit or FP16.

Why does more context need more memory?

The model keeps a KV cache for every token it is considering. Long documents or many users at once multiply it.

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.