TrenchOps 🐎

Insights from the tech trenches

Cover Image

We Abolished Programming Languages - Then We Reinvented Them - In Prose

SupportTrenches Stories & Case Studies 8 minutes

German is objectively denser than English. It still makes your prompts worse. And the reason tells you exactly where this is all heading.

A while ago, someone asked me, with the sincere enthusiasm of a person who had just had a Big Idea: "German is so precise. Wouldn't prompting in German be more token-efficient?"

As a German, I wanted this to be true so badly. Finally, a competitive advantage from the language that gave the world Schadenfreude, Kummerspeck and Weltschmerz.

It is not true. But the reason it fails is more interesting than the idea itself.

The German hypothesis is linguistically sound

German really does pack more meaning per word. Compound nouns fuse concepts into single unambiguous units. Datenbankverbindungspoolverwaltung is one word, one concept, and the structure tells you precisely what belongs to what. The English equivalent, "database connection pool manager", are four words that leave you guessing whether the manager manages the pool, the connections, or the database. Add a case system that encodes relationships without prepositions, plus a vocabulary that prefers specific terms over vague umbrellas, and in theory you get fewer words for the same instruction.

Fewer words, fewer tokens, cheaper and sharper prompts. Beautiful theory.

The tokenizer has other plans

Tokenizers were trained mostly on English. So that gorgeous compound noun does not arrive as one dense unit. It arrives shredded: Dat / enbank / verbind / ungs / pool / verwalt / ung. Six-ish tokens of subword confetti, none of which mean anything, versus two or three clean tokens for the English phrase. You wrote one word and got billed by the syllable.

Then the training imbalance piles on. If English is over half the corpus and German is a rounding error, the model's internal reasoning about technical concepts is anchored in English whether you like it or not. Prompting in German means the model quietly translates, reasons, and translates back - a border crossing at every step, and customs take a cut each time.

And the third problem is the one nobody wants to hear: code is English. Function names, library names, error messages, stack traces, documentation. You can write your specification in flawless German, but the moment the model has to reason about ConnectionPoolExhaustedException, it is back in English anyway. You have not added precision. You have added a translation layer with no test coverage.

The intuition was right, though. Just aimed at the wrong target. The problem is not which natural language. The problem is natural language.

Ambiguity is not a bug in human speech. It is a feature. It is how diplomacy, poetry, flirting and performance reviews work. It is catastrophic in a specification.

The circle nobody expected to close this fast

Look at the trajectory of talking to computers:

The great promise was that you would never need a programming language again. What practitioners actually shipping production systems have converged on is... a programming language. With a nondeterministic runtime, no compiler, and no error messages - just confident nonsense delivered in a friendly tone.

This was inevitable and the mechanism is simple. "Add these two numbers and return just the result" works 98% of the time. In a demo, 98% is magic. In production, 98% is an incident channel. So you bolt on constraints: "return ONLY an integer." "No commentary." "Output valid JSON, nothing else." Each of those sentences exists to close exactly one observed failure mode.

That is not prose. That is a line of code, written in prose, badly.

Which is precisely why typed prompting languages and grammar-constrained generation tools exist - BAML with function signatures, Guidance and Outlines constraining output to a grammar or state machine, LMQL wrapping model calls in query syntax. Every one of them is an honest admission: natural language was a wonderful demo and a poor interface for reliable systems.

What genuinely changed

One thing did change, and it matters. Traditional code specifies HOW - the algorithm, step by step. Structured prompts specify WHAT - the desired outcome plus constraints. That is declarative. Closer to SQL than to C.

So the real evolution reads: imperative (write the HOW) β†’ declarative (write the WHAT) β†’ natural language (vaguely gesture at the WHAT while hoping) β†’ structured prompts (declare the WHAT precisely, with types).

We took a detour through gesturing. It was fun. We are back.

The number that should bother you

Take a rambling 2000-token prompt. Strip out the politeness, the restatements, the three contradictory versions of the same requirement, the "please be thorough and think carefully." What remains as actual decision-relevant content? Often under a hundred tokens.

A 20:1 ratio between what you pay for and what carries information does not mean you need a bigger model. It means the specification is broken. A tight 200-token structured prompt on a cheap model will frequently beat the 2000-token ramble on a frontier model, and cost a fraction. Bigger models mostly buy you tolerance for your own vagueness. Expensive tolerance.

And here is the part relevant to anyone running a support or ops team: the people on your team who are already excellent at this are not the ones with "AI" in their job title. They are the ones who write repro steps a stranger can follow. The ones whose escalations never come back with "can you clarify?" The ones whose knowledge base articles work for someone who has never seen the product. That is the same skill, and you have been undervaluing it for years while it sat two desks over. 🐎

πŸ‘‰ Prompt in the language your tooling, errors and training data live in. For technical work, that is English. Linguistic elegance loses to tokenizer reality.

πŸ‘‰ Density is not precision. Precision is structure: schemas, types, enumerated constraints, explicit failure modes.

πŸ‘‰ Every constraint you add to a prompt is a bug fix. Treat it like one - version it, comment why it exists, and never delete it without knowing which failure it prevented.

πŸ‘‰ Measure the ratio of tokens paid to decisions actually specified. High ratio means fix the spec, not the model.

πŸ‘‰ Declarative beats imperative here. Describe the outcome and its boundaries, not the reasoning steps. Then constrain the output format hard.

πŸ‘‰ Your best "context engineers" are your clearest writers. Find them in support and QA, not in the job postings.

"Context engineering" is just programming with extra steps and worse error messages. The skill was never obsolete - it only changed uniforms.

πŸ‘‡ What is the one constraint you had to add to a prompt after production taught you a lesson the hard way?

This article is also available in German.


AISoftware EngineeringTechnologyDocumentationCommunicationProductivityPrompt Engineering

0 comment(s)

No comments yet. Be the first to comment.

Leave a comment

0 / 1000