Training is when a model learns; inference is when you use what it learned. Asking a chatbot a question, generating an image or transcribing audio are all inference.
Inference needs far less memory and compute than training. Usually the model just has to fit in GPU memory; speed depends mostly on the GPU’s memory bandwidth.
Example
Running Llama 3.1 8B to answer visitors on your website is inference.
বাংলায়: ইনফারেন্স (মডেল চালানো) — শেখা হয়ে যাওয়া মডেলকে দিয়ে কাজ করানো — প্রশ্ন দিলে উত্তর, লেখা দিলে ছবি।