LLMs do not read letters or words directly; they split text into tokens. In English a token is roughly four characters, or three-quarters of a word.
Most models need many more tokens for Bangla than for the same meaning in English — up to six times as many with some tokenizers. API bills and model limits are counted in tokens, so Bangla costs more.
“আমি বাংলায় কথা বলি।” is 27 tokens with GPT-4’s tokenizer, while “I speak Bangla.” is just 5.
বাংলায়: টোকেন — LLM লেখাকে যে ছোট ছোট টুকরোয় ভাগ করে পড়ে। দাম আর সীমা — দুটোই টোকেনে মাপা হয়।