VRAM is the GPU’s fast on-board memory. To run a model, all its parameters (weights) must sit in VRAM, plus the conversation history (the KV cache) and some working space.
Rule of thumb: at FP16 each billion parameters needs about 2 GB; at 4-bit about 0.56 GB. So an 8B model needs roughly 16–18 GB at FP16 and about 6 GB at 4-bit.
Example
A 24 GB L4 runs Llama 3.1 8B at full quality; a 70B model needs an 80 GB A100/H100 and 4-bit.
বাংলায়: VRAM (GPU মেমরি) — GPU-র নিজস্ব মেমরি। মডেলটি পুরোটা এখানে না ধরলে চালানো যায় না।