Skip to content
View original post on X: Prime Intellect· 20/100AI score20/100

Prime Intellect: DEP8 cuts prefix-cache pressure versus TEP8 on same GPUs

AISummary

Prime Intellect reports that DEP8 provides about 5x the prefix-cache capacity of TEP8 on the same GPUs. The post argues that fast KV retrieval alone does not ensure fast first tokens, since cached KV often sat ready while requests waited to join a batch. Halving the prefill budget reduced median queue wait time and time to first token (TTFT).

Post on XView on X
@PrimeIntellect

A reply · the post it answers

Prefill: time to first token

DEP8 gave us ~5x the prefix-cache capacity of TEP8 on the same GPUs.

Fast KV retrieval doesn’t guarantee a fast first token. We also need to reduce the time requests spend waiting to run after their KV is ready.

Cached KV was often ready before the request could join a batch. Halving the prefill budget cut median queue wait time and TTFT.

Source: Prime Intellect · x.comPublished · added here