← AI glossary

VRAM (GPU memory)

The GPU’s own memory. If a model does not fit in it, it cannot run.

VRAM is the GPU’s fast on-board memory. To run a model, all its parameters (weights) must sit in VRAM, plus the conversation history (the KV cache) and some working space.

Rule of thumb: at FP16 each billion parameters needs about 2 GB; at 4-bit about 0.56 GB. So an 8B model needs roughly 16–18 GB at FP16 and about 6 GB at 4-bit.

Example

A 24 GB L4 runs Llama 3.1 8B at full quality; a 70B model needs an 80 GB A100/H100 and 4-bit.

বাংলায়: VRAM (GPU মেমরি) — GPU-র নিজস্ব মেমরি। মডেলটি পুরোটা এখানে না ধরলে চালানো যায় না।

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.