2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns only the last 128. In prefill, layers 21–39 run only on each request's last 128 tokens. With CUDA graphs, prefill compute drops 30–40%.
2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns onl...
AISummary
2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns only the last 128. In prefill, layers 21–39 run only on each request's last 128 tokens. With CUDA graphs, prefill compute drops 30–40%.
Source: vLLM · x.com