AI Grimoire

Inference / serving

Serving

A served model spends its time on a stream of requests of wildly different lengths. Throughput is decided by the scheduler — by whether a finished sequence’s slot can be refilled without waiting for the rest of its batch.

1 entry.

Entries

04.05.1
Continuous Batchingstandard
Iteration-level scheduling; requests join and leave the batch mid-flight.
O(1)