Compression as via negativa: the cheapest token is the one you never send
Take any coding agent session and open the raw request log. Not the pretty transcript - the actual payload. What you find is a JSON search result with 100 hits where 3 mattered, a log file where the interesting line is buried in 800 lines of "INFO: still fine," and a directory listing of a repo the model already walked twice.
Then apply the uncomfortable ratio: cost of the finished output versus the raw inputs actually required to produce it. The model needed maybe 2,000 tokens of real signal. You paid for 55,000. In manufacturing terms, that is a part made of solid gold to hold a plastic clip in place.
I spent a few weeks running everything - coding, research, boring text work - through a compression proxy called Headroom, and the interesting part was not the money. It was watching exactly how much of what I "sent" to the model was never information in the first place.
What it actually does
It sits between your agent and the API and squeezes everything the model reads - tool outputs, logs, retrieved chunks, conversation history - before it goes upstream. Compression happens locally, nothing gets shipped off to a third party to be shrunk. The model can pull the original back if it needs it.
The three words that reveal exactly how little we understand the machines we now depend on
The funniest prompt in modern computing is three words long, and it appears in production systems at companies with actual revenue.
"Don't make mistakes."
Sometimes with emphasis. DO NOT MAKE MISTAKES. Sometimes with a threat attached, as if the model has a family. Sometimes with a promise of a tip, which is my favorite genre of magical thinking - bribing a probability distribution.
Sit with the assumption for a second. Telling a system not to make mistakes implies the system was previously choosing to make them. Which means one of three things must be true: it was trained to err, it was instructed to err (perhaps by a vendor who bills per token - a delightfully paranoid theory that dies the moment you notice open-weight models behave identically), or errors are an opt-in feature that ships enabled by default.
None of it survives ten seconds of contact with how the thing actually works. A language model does not have a laziness dial. It has no intent to be sloppy, because it has no intent at all. It produces the statistically plausible continuation of your text. "Don't make mistakes" is not an instruction. It is a mood. It shifts the output slightly toward the register of text that appears near careful, hedged, authoritative-sounding language in the training data - which is why the response often sounds more confident while being exactly as wrong. You did not reduce the error rate. You upgraded the packaging.
The internet already ran this experiment. We just didn't like the results.
Do you remember dial-up? Paying per minute to access the sum of human knowledge, and using those precious minutes to download a 400x300 JPEG of a cat in a shoebox?
Then the meter went away. Flat rate. Then mobile. Then free WiFi in the average bakery. Suddenly every human on the planet had a portal to every library, every lecture, every tutorial, every trade skill ever documented. Universities put their entire curriculum online for free. MIT did it. Harvard did it. Nobody had to ask permission anymore.
And what happened?
TikTok happened. Instagram happened. LinkedIn happened, which is basically Instagram for people who own a blazer. Millions of educational videos exist on YouTube, and the most watched content is people reacting to other people reacting to video games.
This is not a rant. I like cat pictures. I have watched a man restore a rusty axe for 22 minutes and felt genuine peace afterwards. The observation is colder than judgment: unlimited access to knowledge did not redistribute outcomes. The people who were competent before the internet were mostly competent after it. The people who were drifting kept drifting, just with better graphics.
For fifteen years, "the cloud is infinite" was the most reliable lie in enterprise IT. Not a malicious lie. A useful one. It let architects design for peak load without capacity planning, let CFOs treat compute as opex, and let me tell customers "just scale out" with a straight face.
Then customers in North America and the UK started filing tickets that read like they came from 2004. "Cannot provision." "Cannot reschedule." "Try another region." Across all three hyperscalers, in specific regions, at specific instance families. AWS reportedly told its own engineers to conserve compute "however they can" - internal teams waiting days for CPUs, because customer workloads come first. Which is the correct priority and also a sentence that should not exist in a business built on the promise of on-demand.
Ten years ago I would have laughed at anyone predicting this. The whole pitch was that Amazon, Google and Microsoft had so much spare iron that your workload was a rounding error. And they did. They just sold the rounding error to language models.
The number that shouldn't exist
Here is the part that should make every architect uncomfortable. Cast AI's 2026 telemetry across 23,000+ production clusters puts average GPU utilization at about 5%. On AKS, 2%. On EKS, 5%. On GKE, 6%.
Why I moved back to SUSE after two decades - and why you should re-evaluate your stack while it still works
Somewhere around the year 2000, I bought a computer magazine because it had a CD glued to the cover. SUSE Linux 6.4. ReiserFS. I got as far as "select a mount point for root" and stopped, because I had no idea what a mount point was, and no realistic way to find out. No search engine within reach. No forum. Just me, a beige tower, and a dialog box asking me a question in a language I didn't speak.
I abandoned the installation. My first contact with Linux ended in surrender at the partitioning screen, which I suspect is the most common origin story in our industry.
A few months later, SUSE 7 came out and I ordered the boxed set. Seven CDs. A printed manual thick enough to stop a small projectile. A t-shirt. A sticker of a chameleon. And most importantly: no license key scrawled on a disc in marker by someone's cousin. I installed an operating system that nobody asked me to prove I deserved. That smell of freedom stayed with me longer than the T-shirt did.
German is objectively denser than English. It still makes your prompts worse. And the reason tells you exactly where this is all heading.
A while ago, someone asked me, with the sincere enthusiasm of a person who had just had a Big Idea: "German is so precise. Wouldn't prompting in German be more token-efficient?"
As a German, I wanted this to be true so badly. Finally, a competitive advantage from the language that gave the world Schadenfreude, Kummerspeck and Weltschmerz.
It is not true. But the reason it fails is more interesting than the idea itself.
The German hypothesis is linguistically sound
German really does pack more meaning per word. Compound nouns fuse concepts into single unambiguous units. Datenbankverbindungspoolverwaltung is one word, one concept, and the structure tells you precisely what belongs to what. The English equivalent, "database connection pool manager", are four words that leave you guessing whether the manager manages the pool, the connections, or the database. Add a case system that encodes relationships without prepositions, plus a vocabulary that prefers specific terms over vague umbrellas, and in theory you get fewer words for the same instruction.
Fewer words, fewer tokens, cheaper and sharper prompts. Beautiful theory.
A hardware wallet built keys from a random number generator it never once called. For five years. The source code was public the entire time.
In a 25-minute window, someone swept roughly 594 bitcoin - about $38 million - out of some 500 unrelated wallets. No phishing. No malware. No supply chain tampering. No 5$ wrench.
Every one of those wallets was protected by a device sold on a single, admirable promise: keys generated offline, on isolated hardware, using a true hardware random number generator, Bitcoin-only to shrink the attack surface, no cloud, no companion app, and open-source firmware you can verify yourself.
The hardware RNG was on the board. Powered. Functional.
For roughly 1,900 days, the firmware never asked it for a number.
The bug is one character wide
The vendor did something entirely sensible: they switched off MicroPython's built-in random path with a build flag, because they had written their own wrapper around the microcontroller's true RNG. Textbook embedded hygiene.
Then a cryptographic support library checked that flag with #ifndef - which asks "does this macro exist?" rather than "is it set to something other than zero?"
The macro existed. Its value was zero. So the library concluded hardware entropy was available and happily bound itself to MicroPython's random function. MicroPython, having read the same macro correctly, had compiled a non-cryptographic software fallback instead of the hardware peripheral.
When your entire career is a YAML file, don't be surprised when someone auto-generates it
A while back, I made an offhand remark somewhere about Kubernetes being overkill for a particular use case. Nothing radical. Just pointing out that a single application serving a few thousand users probably doesn't need a container orchestration platform designed to run Google's entire production fleet.
Within minutes, the clergy arrived.
"C'mon bro. Just pipe your Jsonnet through the Kluster Konfig Kompiler, deploy the sidecar injector via a mutating webhook, add the ChaosMonkeyMesh operator for resilience testing, sync your GitOps state through FluxKapacitor, and tail the logs from the Kloud-Native Kombined Kockpit. It's literally five steps. Why are you so afraid of YAML?"
I'm not afraid of YAML. I'm afraid of people who think memorizing a toolchain is the same as understanding infrastructure.
Here's the thing about Kubernetes. It's a genuinely good piece of technology. Google donated it to the world, the CNCF nurtured it, and it solved a real problem - orchestrating containers at scale across distributed systems. If you're running hundreds of microservices across multiple regions with complex networking, auto-scaling requirements, and zero-downtime deployments, Kubernetes is probably the right call.
But somewhere along the way, "right tool for certain jobs" became "the only tool for every job." And an entire professional subculture emerged around that confusion.
And that silence is the most accurate forecast we have for AI content
Remember when installing a free screensaver was an act of war?
Twenty years ago, the internet had a moral panic with its own industry attached: spyware. Purple gorillas reading your email. Browser toolbars reproducing like rabbits. People wrote furious forum posts. Lawmakers held hearings. Anti-spyware was an entire software category with boxed products and annual renewals.
Today your TV, your car, your phone, and - God help us - your fridge all phone home continuously, and the strongest public reaction is a mildly annoyed click on "Accept all."
What happened? The users didn't win. The word lost.
Spyware became "telemetry." Telemetry became "diagnostics." Diagnostics became "personalized experiences." The practice never changed - the vocabulary did. A controversial technology doesn't need to be accepted to win. It just needs its name retired. Once nobody can say the old word without sounding like a crank in a tinfoil hat, the practice has become infrastructure.
Now watch the same movie on 4x speed.
Stage one: ridicule. Hands with seven fingers. A Hollywood star unhinging his jaw over a plate of spaghetti. That was the golden age - spotting AI was a party trick, and some of us pattern-matchers enjoyed it the way birdwatchers enjoy rare warblers.
Some years back, a company that sold iced tea renamed itself "Long Blockchain Corp." The stock jumped almost 300% in a single day. They had no blockchain. They had no blockchain engineers. They had, to be fair, very drinkable iced tea. The company was later delisted, the SEC revoked its registration, and the whole episode ended with insider trading charges around the announcement.
Read that again: the announcement moved the price 300%. Not a product. Not a prototype. A word.
That number shouldn't exist. But it does, and it explains the entire era better than any whitepaper ever did.
The meeting I still think about
Mid-hype, a client asked me to evaluate putting their support knowledge base "on the blockchain." I asked my standard opening question, the one that has ended more projects than any budget cut: which problem does this solve that a database with an audit log doesn't?
Long silence. Then, honestly, to their credit: "Our biggest competitor announced a blockchain initiative last quarter."
There it is. The whole mechanism, in one sentence.
Nobody in that room wanted blockchain. Nobody had derived it from a problem, worked backwards from a customer pain, or hit a wall that only a distributed ledger could break through. They wanted it because someone else visibly wanted it. The desire was borrowed. The entire enterprise blockchain wave was one giant chain of companies watching each other's press releases and concluding "they must know something we don't" - while the other side of the chain was thinking the exact same thing about them.