vLLM bounded SWA replay cuts prefill compute 30–40% after prefix hits
AIvLLM's bounded sliding-window attention (SWA) replay reruns only each request's last 128 tokens after a prefix hit, instead of replaying 40 × 128 tokens. With CUDA graphs, prefill compute drops 30–40% for layers 21–39.







