← AI glossary

Quantization

Storing parameters in fewer bits to shrink a model — 4-bit needs about a quarter of the memory.

Normally each parameter takes 16 bits (FP16). Quantization stores them in 8 or 4 bits, so the model is smaller and faster with only a small loss of quality.

GGUF, AWQ and GPTQ are popular 4-bit formats. For chatbots 4-bit is often enough; for precise work 8-bit or FP16 is safer.

Example

Llama 3.1 70B is ~140 GB at FP16 but ~40 GB at 4-bit — it fits on one 80 GB A100.

বাংলায়: কোয়ান্টাইজেশন — প্যারামিটারগুলো কম বিটে লিখে মডেলকে ছোট করা — ৪-বিটে মেমরি প্রায় চার ভাগের এক ভাগ।

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.