Can a model tell which region a Bangla sentence comes from? This notebook trains a text classifier on the BD-Dialect dataset (CC BY 4.0) and compares standard Bangla BERT models side by side.
What you do
- Load and inspect the dialect data; map the columns once if names differ.
- Fine-tune a classifier and read the per-dialect scores.
- Compare models and their licences.
Good to know
- The well-known BanglaBERT (csebuetnlp) is CC BY-NC-SA 4.0 — non-commercial only. The notebook defaults to an MIT-licensed model so you can build commercial work, and shows how to switch if your use is research.
- A small GPU (24 GB) is plenty.