throughput

TPOT

Time per output token. After the first word arrives, this is the metronome: how fast every following word appears, one after another, until the answer is done.

TPOT is the steady-state clock of decoding — memory bandwidth divided by model size, mostly. It says nothing about how long you waited to start, and everything about how the rest of the answer reads: streaming at reading speed, or arriving as text you scroll.

Serving engineers trade it against TTFT constantly. Batch more requests and each one’s TPOT slows, but the fleet gets cheaper. Speculative decoding, better kernels, quantization — they all show up here first: tokens per second, per stream, per user.

The first token gets the glory. This one does the work.

ttft.siThe wait before the first word — the other half of the tradeoff. kvcache.siThe memory every one of these tokens reads from.