Overview

As conversations grow, they can exceed the model’s context window. Loop’s compaction system automatically manages this by summarizing older parts of the conversation while keeping recent context intact.

How It Works

  1. Loop tracks token usage across the conversation
  2. When the context approaches the model’s limit, compaction triggers
  3. Older messages are summarized while recent messages are preserved
  4. The compacted conversation continues seamlessly

Configuration

Configure compaction in ~/.loop/agent/settings.json:

Manual Compaction

Trigger compaction manually with a slash command:
You can also provide guidance for the compaction:
The guidance helps the summarizer prioritize what to preserve in the summary.

Token Tracking

View your current token usage:
This shows:
  • Input tokens — Total tokens sent to the model
  • Output tokens — Total tokens generated
  • Cache reads/writes — Tokens served from and written to cache
  • Cache waste — Tokens that were cached but never re-read
  • Estimated cost — Approximate session cost

Strategies

For extended sessions, compaction keeps the conversation productive. The agent retains awareness of what was discussed while freeing context for new work.
Use /compact "focus on X" to guide what information is preserved. This is useful when you want to shift focus within a long session.
If compaction isn’t enough, use /new to start a fresh session. You can always /resume the previous session later.
A session_before_compact hook can supply the summary, or cancel compaction. Otherwise Loop uses a built-in summary of the cut messages and keeps the recent tail. Set compaction.enabled to false to turn compaction off.