Your AI Agent Reads 50,000 Tokens to Learn One Thing
Compression as via negativa: the cheapest token is the one you never send
Take any coding agent session and open the raw request log. Not the pretty transcript - the actual payload. What you find is a JSON search result with 100 hits where 3 mattered, a log file where the interesting line is buried in 800 lines of "INFO: still fine," and a directory listing of a repo the model already walked twice.
Then apply the uncomfortable ratio: cost of the finished output versus the raw inputs actually required to produce it. The model needed maybe 2,000 tokens of real signal. You paid for 55,000. In manufacturing terms, that is a part made of solid gold to hold a plastic clip in place.
I spent a few weeks running everything - coding, research, boring text work - through a compression proxy called Headroom, and the interesting part was not the money. It was watching exactly how much of what I "sent" to the model was never information in the first place.
What it actually does
It sits between your agent and the API and squeezes everything the model reads - tool outputs, logs, retrieved chunks, conversation history - before it goes upstream. Compression happens locally, nothing gets shipped off to a third party to be shrunk. The model can pull the original back if it needs it.
The project's own benchmarks show roughly 20% on code search, 40-60% on incident debugging and codebase exploration, and the big numbers (86%+) on JSON arrays, which makes sense: JSON is mostly punctuation and repeated key names, a format designed for parsers and billed as if it were prose. Structured logs compress hard too. Source code mostly passes through untouched, which is the right default - you do not want your agent reasoning about a lossy version of the file it is editing.
Honest caveat, because I would rather you trust me in six months: the dashboard counts tokens it compressed, not tokens your provider stopped charging you for. Those are related but not identical numbers, and some users report the delta being smaller than advertised once retrieval round-trips are included. Treat the dashboard as a direction indicator, not an invoice.
Which brings up the part nobody markets: if you are on a flat subscription plan, you save exactly zero euros/dollars. What you save is runway. Fewer tokens consumed means you travel further before slamming into the 5-hour and weekly usage windows. For anyone who has had a long refactor cut off mid-thought by a rate limit, that is worth more than a discount.
Setup on Linux, in the time it takes to make coffee
Install and verify:
uv tool install --python 3.13 "headroom-ai[all]"command -v headroomheadroom doctor
You want "Proxy" showing a green checkmark. If it does not, stop here and fix that first - do not proceed on hope.
Optional but sensible, run it as a user daemon so it is simply always there:
# ~/.config/systemd/user/headroom.service[Unit]Description=Headroom compression proxy
[Service]ExecStart=%h/.local/bin/headroom proxy --port 8787Restart=on-failure
[Install]WantedBy=default.target
systemctl --user enable --now headroomss -lntp | grep 8787
That second command is not cosmetic. Confirm it is bound to 127.0.0.1 and not 0.0.0.0. An unauthenticated proxy that sees every prompt you write, listening on all interfaces, is not a productivity tool. It is a gift to whoever else is on your network.
Then point your client at it, in ~/.claude/settings.json:
{Β "env": {Β "ANTHROPIC_BASE_URL": "http://127.0.0.1:8787"Β }}
Restart the CLI and your IDE - VS Code or VSCodium will not pick up the new base URL otherwise. Open http://127.0.0.1:8787/dashboard and check that "Request Health" and "Live Activity" show non-zero numbers. That is the whole installation. The rest is watching a counter go up while you do the work you were going to do anyway.
When it breaks, and it will
It is a young, fast-moving project. Rarely, requests hang or die. The triage order, learned the boring way:
π Stop the current model session, restart the tool or IDE, tell the agent "continue." Fixes most of it.
π Still failing? systemctl --user restart headroom.
π Still failing? It is probably not your machine. Check status.claude.com before you debug anything else - degraded upstream service looks exactly like a broken proxy from where you are sitting, and I have wasted honest minutes proving that.
π Genuinely a regression in a new release? Pin backwards: uv tool install --python 3.13 "headroom-ai[all]==0.37.0" --force. Fast-moving projects reward people who know how to step back one version instead of filing an issue and waiting.
That last point is the real skill, by the way. Not the tool - the habit of having a rollback path before you need one.
The part that generalizes
Compression proxies are a workaround. The underlying problem is that we hand agents firehoses and call it context. We pipe in complete API responses, full log files and entire directory trees because it is easier than deciding what matters, then pay per token for the privilege of making a language model do our filtering at premium rates.
The best part is no part. The cheapest token is the one you never send. A tool that compresses your junk is strictly better than not having it - but a grep with a sane filter, a log level that is not DEBUG in production, and an API response that returns fields instead of everything would have gotten you most of the way there without any middleware at all.
π What is the dumbest thing your agent has ever been made to read in full? Mine spent real money ingesting a 400-line INFO log to find one timestamp.
TrenchOps π
0 comment(s)
No comments yet. Be the first to comment.
Leave a comment