← AI glossary

QLoRA

LoRA on top of a 4-bit compressed base model — fine-tuning with the least memory.

QLoRA keeps the base model quantized to 4-bit and trains LoRA adapters on top, cutting memory to roughly a third or a quarter of plain LoRA.

The trade-off is speed: each step is about 20–30 % slower. It is the most practical way to fine-tune a large model on a smaller GPU.

Example

A QLoRA fine-tune of Qwen2.5 32B fits on a 48 GB L40S.

বাংলায়: QLoRA — মূল মডেলকে ৪-বিটে ছোট করে তার ওপর LoRA — সবচেয়ে কম মেমরিতে ফাইন-টিউনিং।

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.