Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Just to avoid some confusion: auto-compact does summarize. When it runs, your whole conversation is replaced by a short summary. It doesn't keep the last 1M tokens around. The ~967K is only when it runs on 1M models. Compacting itself is just one request that reads about as many tokens as the message you sent right before it, and it's mostly cache reads if your cache is still warm. What @rohit3a found is that waiting until 967K means every message before that point uses a lot of tokens,…
holy fuck. so auto compact basically loads last 1M tokens into context by default on claude? and here i thought they were summarising and reducing the context loads. no wonder people used to burn anywhere from 2%-15% 5h limits simply for auto compacting their sessions. https://t.co/RzH6YpvNUe