← All free tools

How much VRAM does Mixtral 8x7B need?

Mixtral 8x7B needs about 95.4 GB of GPU memory at FP16 and about 27.9 GB at 4-bit. Pick a precision below to see which ComputeBD AI Computer fits and what it costs per hour in taka.

What do you want to do?

You need about 95.4 GB of GPU memory to run Mixtral 8x7B at FP16 / BF16.

  • Model weights 87 GB
  • Conversation memory (KV cache) 0.5 GB
  • Runtime & working space 8 GB

NVIDIA H200 is the most affordable AI Computer that fits — ৳1,390/hour, billed by the minute.

AI ComputerFits?Speed (1 user)Price
NVIDIA L424 GB Too small
≈ 7 tokens/s ৳147/hr
NVIDIA A1024 GB Too small
≈ 14 tokens/s ৳196/hr
NVIDIA L40S48 GB Too small
≈ 20 tokens/s ৳343/hr
NVIDIA A10080 GB Too small
≈ 47 tokens/s ৳638/hr
NVIDIA H10080 GB Too small
≈ 78 tokens/s ৳1,047/hr
NVIDIA H200141 GB Fits
≈ 112 tokens/s ৳1,390/hr
NVIDIA B200180 GB Fits
≈ 186 tokens/s ৳1,652/hr
NVIDIA B300288 GB Fits
≈ 186 tokens/s ৳1,880/hr

Mixtral 8x7B at each precision

PrecisionGPU memory neededCheapest ComputeBD AI Computer
FP32full precision — rarely needed189.4 GBNVIDIA B300 · ৳1,880/hr
FP16 / BF16standard — original quality95.4 GBNVIDIA H200 · ৳1,390/hr
8-bitalmost no quality loss51.4 GBNVIDIA A100 · ৳638/hr
4-bit (GGUF / AWQ / GPTQ)small quality loss, ~4× less memory27.9 GBNVIDIA L40S · ৳343/hr
Made by
Mistral AI
Parameters
46.7B (12.9B active per token)
Max context
32,768 tokens
LoRA / QLoRA fine-tune
≈ 102.4 GB / 39.9 GB
Hugging Face
mistralai/Mixtral-8x7B-Instruct-v0.1

How the estimate works

These are good planning estimates, not guarantees: the real figure depends on the software (vLLM, llama.cpp, Transformers…), settings and your data. Leave ~10 % spare.

VRAM needed, model by model

Frequently asked questions

How much VRAM does Mixtral 8x7B need?

About 95.4 GB at FP16 (original quality) and about 27.9 GB at 4-bit, for one user with a 4k context.

Which GPU can run Mixtral 8x7B?

At full quality it fits on an NVIDIA H200 (141 GB), ৳1,390/hour on ComputeBD.

How much VRAM to fine-tune Mixtral 8x7B?

About 102.4 GB with LoRA and 39.9 GB with QLoRA (sequence length 2,048, batch 1, gradient checkpointing). Estimate the time and cost with the fine-tuning cost estimator.

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.