Skip to content
Read the original: vLLM· vllm_project·Published· YesterdayAI score28

2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns only the last 128. In prefill, layers 21–39 run only on each request's last 128 tokens. With CUDA graphs, prefill compute drops 30–40%.

2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns onl...

AISummary

2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns only the last 128. In prefill, layers 21–39 run only on each request's last 128 tokens. With CUDA graphs, prefill compute drops 30–40%.

Read the original x.com

Source: vLLM · x.com