Just to avoid some confusion: auto-compact does summarize. When it runs, your whole conversation is replaced by a short summary. It doesn't keep the last 1M tokens around.
The ~967K is only when it runs on 1M models. Compacting itself is just one request that reads about as many tokens as the message you sent right before it, and it's mostly cache reads if your cache is still warm.
What @rohit3a found is that waiting until 967K means every message before that point uses a lot of tokens, because each one includes your whole conversation so far. Most of it is read from the cache, which is much cheaper (so this is usually fine as long as your cache is warm), but it still counts toward your usage and adds up as your conversation grows.
/autocompact 400k makes it summarize at 400K instead, so your messages never get that big.
