Read the original: PyTorch Blog· Ishan Dhanani, Jamie Li, Karen Chung·Published· 9h agoPickAI score62
NVIDIA Dynamo adds session-level IDs to route and cache agentic inference
Session-Aware Agentic Inference with NVIDIA Dynamo
AISummary
NVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.
AIWhy it matters
The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.
Source: PyTorch Blog · pytorch.org