The context window is the limit of a model’s working memory. Llama 3.1 allows 128,000 tokens — about a short book.
Longer contexts use more GPU memory (the KV cache), because information for every token must be kept.
Example
For Llama 3.1 8B the KV cache is about 0.5 GB at 4,000 tokens and about 16 GB at 128,000.
বাংলায়: কনটেক্সট উইন্ডো — মডেল একবারে কত টোকেন মনে রাখতে পারে — প্রশ্ন, ডকুমেন্ট আর উত্তর মিলিয়ে।