QLoRA keeps the base model quantized to 4-bit and trains LoRA adapters on top, cutting memory to roughly a third or a quarter of plain LoRA.
The trade-off is speed: each step is about 20–30 % slower. It is the most practical way to fine-tune a large model on a smaller GPU.
Example
A QLoRA fine-tune of Qwen2.5 32B fits on a 48 GB L40S.
বাংলায়: QLoRA — মূল মডেলকে ৪-বিটে ছোট করে তার ওপর LoRA — সবচেয়ে কম মেমরিতে ফাইন-টিউনিং।