← AI glossary

Context window

How many tokens a model can consider at once — prompt, documents and answer together.

The context window is the limit of a model’s working memory. Llama 3.1 allows 128,000 tokens — about a short book.

Longer contexts use more GPU memory (the KV cache), because information for every token must be kept.

Example

For Llama 3.1 8B the KV cache is about 0.5 GB at 4,000 tokens and about 16 GB at 128,000.

বাংলায়: কনটেক্সট উইন্ডো — মডেল একবারে কত টোকেন মনে রাখতে পারে — প্রশ্ন, ডকুমেন্ট আর উত্তর মিলিয়ে।

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.