← AI glossary

Inference

Using a trained model to do work — answering a question, generating an image.

Training is when a model learns; inference is when you use what it learned. Asking a chatbot a question, generating an image or transcribing audio are all inference.

Inference needs far less memory and compute than training. Usually the model just has to fit in GPU memory; speed depends mostly on the GPU’s memory bandwidth.

Example

Running Llama 3.1 8B to answer visitors on your website is inference.

বাংলায়: ইনফারেন্স (মডেল চালানো) — শেখা হয়ে যাওয়া মডেলকে দিয়ে কাজ করানো — প্রশ্ন দিলে উত্তর, লেখা দিলে ছবি।

Try it on a real AI Computer

JupyterLab opens in about a minute. Pay in taka with bKash, billed by the minute — stop whenever you like.