Prefill: time to first token
DEP8 gave us ~5x the prefix-cache capacity of TEP8 on the same GPUs.
Fast KV retrieval doesn’t guarantee a fast first token. We also need to reduce the time requests spend waiting to run after their KV is ready.
Cached KV was often ready before the request could join a batch. Halving the prefill budget cut median queue wait time and TTFT.
