TTFT (time-to-first-token) is one of the most-quoted latency numbers in inference.
But does TTFT actually matter, or is it the most overrated metric in the stack?
The real axis isn't "fast." It's interactivity (tok/s/user) vs throughput (tok/s/GPU). (1/3)🧵