Skip to content
View original post on X: Lydia Hallie ✨X· 36/100AI score36/100

Claude Code auto-compact summarizes conversations, not the last 1M tokens

AISummary

Anthropic's Lydia Hallie clarifies that Claude Code's auto-compact replaces the whole conversation with a short summary. On 1M-context models it triggers around 967K tokens, and each message before that point re-reads the full conversation, mostly from cache. Running /autocompact 400k makes compaction trigger at 400K instead.

Post on XView on X
Lydia Hallie ✨Verified on X
@lydiahallie

Just to avoid some confusion: auto-compact does summarize. When it runs, your whole conversation is replaced by a short summary. It doesn't keep the last 1M tokens around.

The ~967K is only when it runs on 1M models. Compacting itself is just one request that reads about as many tokens as the message you sent right before it, and it's mostly cache reads if your cache is still warm.

What @rohit3a found is that waiting until 967K means every message before that point uses a lot of tokens, because each one includes your whole conversation so far. Most of it is read from the cache, which is much cheaper (so this is usually fine as long as your cache is warm), but it still counts toward your usage and adds up as your conversation grows.

/autocompact 400k makes it summarize at 400K instead, so your messages never get that big.

Lxyv@wholyv
holy fuck. so auto compact basically loads last 1M tokens into context by default on claude? and here i thought they were summarising and reducing the context loads. no wonder people used to burn anywhere from 2%-15% 5h limits simply for auto compacting their sessions.
View quoted post on X

Source: Lydia Hallie ✨ · x.comPublished · added here