Fine-tuning takes a ready-made model (such as Llama or Qwen) and trains it a little more on your own examples. It learns your language, tone and subject — legal wording, farming advice, or how your organisation answers questions.
Techniques such as LoRA and QLoRA train small adapters instead of the whole model, so it fits on one GPU and finishes in hours.
Turning Qwen2.5 7B into a farming advisor with 5,000 question–answer pairs — a few hours on an L4.
বাংলায়: ফাইন-টিউনিং — তৈরি মডেলকে নিজের ডেটা দিয়ে আরও শেখানো, যাতে সে আপনার কাজে এক্সপার্ট হয়।