NVIDIA L40S vs NVIDIA L4
Specs, monthly cost in taka and the workloads each GPU suits, side by side.
Tensor figures are NVIDIA's published numbers with sparsity (dense is half). Source: L40S, L4.
The short answer
Pick the NVIDIA L40S if: You generate images or video (FLUX, SDXL), fine-tune 7–13B models with LoRA, or serve a 13B model in FP16 — 48 GB for little more than half the A100's hourly rate.
Pick the NVIDIA L4 if: You want to serve a chatbot or API on an 8B model (Llama 3.1 8B, Qwen 2.5 7B), run Whisper or do QLoRA fine-tunes of small models at the lowest hourly cost.
The NVIDIA L40S costs about 133% more per hour than the NVIDIA L4. If your model fits comfortably on the cheaper GPU and you are not short of time, the cheaper one usually wins. Compare with your own model →
Common questions
Which is faster, the NVIDIA L40S or the NVIDIA L4?
On paper the NVIDIA L40S is faster: 733 TFLOPS vs 242 TFLOPS FP16 tensor (with sparsity). Real-world speed depends on your model, batch size and memory bandwidth.
Which one costs less?
The NVIDIA L4 is ৳147 per hour and the NVIDIA L40S is ৳343, about 133% more. For 40 hours that is ৳5,880 vs ৳13,720.
Which has more memory?
The NVIDIA L40S has 48 GB and the NVIDIA L4 has 24 GB. Roughly, the NVIDIA L40S fits 13B in FP16, about 70B at 4-bit (tight); the NVIDIA L4 fits 8B in FP16, about 30B at 4-bit.
Which should I choose for my project?
Choose the NVIDIA L40S if: You generate images or video (FLUX, SDXL), fine-tune 7–13B models with LoRA, or serve a 13B model in FP16 — 48 GB for little more than half the A100's hourly rate. Choose the NVIDIA L4 if: You want to serve a chatbot or API on an 8B model (Llama 3.1 8B, Qwen 2.5 7B), run Whisper or do QLoRA fine-tunes of small models at the lowest hourly cost.