TrenchOps 🐎

Insights from the tech trenches

Token Optimization

Cover Image

Your AI Agent Reads 50,000 Tokens to Learn One Thing

Compression as via negativa: the cheapest token is the one you never send

Take any coding agent session and open the raw request log. Not the pretty transcript - the actual payload. What you find is a JSON search result with 100 hits where 3 mattered, a log file where the interesting line is buried in 800 lines of "INFO: still fine," and a directory listing of a repo the model already walked twice.

Then apply the uncomfortable ratio: cost of the finished output versus the raw inputs actually required to produce it. The model needed maybe 2,000 tokens of real signal. You paid for 55,000. In manufacturing terms, that is a part made of solid gold to hold a plastic clip in place.

I spent a few weeks running everything - coding, research, boring text work - through a compression proxy called Headroom, and the interesting part was not the money. It was watching exactly how much of what I "sent" to the model was never information in the first place.

What it actually does

It sits between your agent and the API and squeezes everything the model reads - tool outputs, logs, retrieved chunks, conversation history - before it goes upstream. Compression happens locally, nothing gets shipped off to a third party to be shrunk. The model can pull the original back if it needs it.

Read more