Instead of updating after every example, the model looks at a batch and updates once. Bigger batches use the GPU more efficiently but need more memory.
When memory is short, gradient accumulation combines several small batches to act like a large one.
Example
Out of memory on an L4? Drop batch size from 8 to 2 and set gradient accumulation to 4.
বাংলায়: ব্যাচ সাইজ — এক ধাপে মডেল কয়টি উদাহরণ একসাথে দেখে।