Speech-to-text models often stumble on regional Bangla. This notebook fine-tunes Whisper on the FLEURS Bengali set (CC BY 4.0) and then measures word error rate per region, so you can see where it fails — for example Sylheti or Chittagonian speakers — before you collect more data.
What you do
- Fine-tune Whisper with a Bangla-safe text normaliser (the default one removes Bengali vowel signs).
- Split data by sentence so that test sentences never leak into training.
- Compare error rates region by region.
Good to know
- Whisper is Apache-2.0. Regional audio sets vary in licence; the notebook shows how to plug in your own recordings and says to record people only with their written consent.