← All notebooks for this

Bangla dialect classifier (BanglaBERT)

Train a text classifier on the BD-Dialect dataset; compare standard Bangla BERT models and their licences.

Can a model tell which region a Bangla sentence comes from? This notebook trains a text classifier on the BD-Dialect dataset (CC BY 4.0) and compares standard Bangla BERT models side by side.

What you do

  • Load and inspect the dialect data; map the columns once if names differ.
  • Fine-tune a classifier and read the per-dialect scores.
  • Compare models and their licences.

Good to know

  • The well-known BanglaBERT (csebuetnlp) is CC BY-NC-SA 4.0 — non-commercial only. The notebook defaults to an MIT-licensed model so you can build commercial work, and shows how to switch if your use is research.
  • A small GPU (24 GB) is plenty.