TrenchOps 🐎

Insights from the tech trenches

Cover Image

"Don't Make Mistakes"

SupportTrenches Stories & Case Studies 9 minutes

The three words that reveal exactly how little we understand the machines we now depend on

The funniest prompt in modern computing is three words long, and it appears in production systems at companies with actual revenue.

"Don't make mistakes."

Sometimes with emphasis. DO NOT MAKE MISTAKES. Sometimes with a threat attached, as if the model has a family. Sometimes with a promise of a tip, which is my favorite genre of magical thinking - bribing a probability distribution.

Sit with the assumption for a second. Telling a system not to make mistakes implies the system was previously choosing to make them. Which means one of three things must be true: it was trained to err, it was instructed to err (perhaps by a vendor who bills per token - a delightfully paranoid theory that dies the moment you notice open-weight models behave identically), or errors are an opt-in feature that ships enabled by default.

None of it survives ten seconds of contact with how the thing actually works. A language model does not have a laziness dial. It has no intent to be sloppy, because it has no intent at all. It produces the statistically plausible continuation of your text. "Don't make mistakes" is not an instruction. It is a mood. It shifts the output slightly toward the register of text that appears near careful, hedged, authoritative-sounding language in the training data - which is why the response often sounds more confident while being exactly as wrong. You did not reduce the error rate. You upgraded the packaging.

That is the trap. The prompt works on the reader, not on the model.

Five groups, and only one of them is dangerous

After too many conversations on this topic, I've stopped arguing about AI in general. There is no "AI debate." There are five different populations having five different conversations while using the same word.

Notice the shape: the risk curve is not linear with enthusiasm. It peaks in the middle. A knows nothing and is safe because its stakes are low. E knows everything and is safe because it knows where the mines are. C knows just enough to be confident, in a domain where confidence is the actual failure mode.

Why the middle is where the damage happens

Here is the mechanism, and it's worth understanding because it's not really about AI.

These systems are trained, among other things, to produce responses humans rate highly. Humans rate agreement highly. So the models are structurally inclined to agree with you - and recent research on chatbot sycophancy suggests something uncomfortable: an agreeable interlocutor can push even a perfectly rational, evidence-updating reasoner toward increasingly wrong conclusions, simply because every step gets confirmed. You don't need to be gullible to spiral. You just need a partner who never pushes back.

Now hand that partner to Group C - a person whose primary skill is enthusiasm and whose primary output is forwarding. The machine agrees. The human feels validated. The text goes out. Nobody in the loop bears any downside, because the author is technically the model and the model has no reputation, no license to lose, and no performance review.

That's the real problem, and it has nothing to do with token probabilities. It's an accountability vacuum wearing a productivity costume. In every functioning system, the person who signs bears the cost of being wrong. The meat proxy signs and bears nothing. Remove skin from the game and quality becomes optional, immediately.

Second-order effect, which I find more worrying than any hallucination: junior people used to build judgment by producing bad work and having it corrected. That's how expertise regenerates - through the expensive, humiliating, irreplaceable loop of being wrong in front of someone senior. If the first draft is always machine-perfect-ish, that loop never runs. There are researchers now describing this as a commons problem: the pool of professional expertise everyone draws from doesn't refill itself if nobody pays the cost of learning anymore. Every individual has a rational incentive to skip the struggle. Collectively, we run out of people capable of noticing the machine is wrong.

Which is exactly the population you need most.

Moving people from C to D

You cannot train this with a policy document. I've watched organizations try. The AI Usage Guidelines PDF gets summarized by AI and forwarded by a meat proxy, which at least demonstrates commitment to the bit.

What works is much smaller and much less pleasant:

πŸ‘‰ Restore the signature. Whoever sends it, owns it. Not "the AI got that wrong" - you got that wrong, in front of the customer. One instance of this landing properly does more than any workshop.

πŸ‘‰ Make the cost visible. Ask what verifying the output cost versus what producing it manually would have cost. Sometimes the honest answer is that a 40-second generation created 25 minutes of fact-checking. That ratio is the whole conversation. If nobody is measuring it, the tool is a hobby, not a process.

πŸ‘‰ Ban the incantations. "Don't make mistakes," threats, imaginary tips. Replace with the things that actually reduce error: give it the source material (e.g. via RAG), constrain the output format, ask for citations you then click, and split hard tasks into checkable steps. Grounding beats begging.

πŸ‘‰ Protect the learning loop. Juniors still write the first draft themselves sometimes. Not for output quality - for the muscle. You are not paying for today's document, you are paying for someone who can still spot a plausible lie in three years.

πŸ‘‰ Reward the person who says "I checked." Verification is invisible work. Invisible work dies unless leadership names it out loud. The quiet power horse 🐎 who caught the fabricated regulation before it reached the client just saved you more than the entire tooling budget.

Group C is not stupid. That's the whole point of the label. They're often the most motivated people you have, pointed slightly wrong. The distance from C to D is not intelligence, it's one habit: read it before you send it.

Nobody who understands how these systems work has ever typed "don't make mistakes." Not because they're above it - because they know exactly which words are load-bearing, and politeness toward a probability distribution isn't one of them (although I know some people who use the word "please" a lot with their LLMs... just in case).

πŸ‘‡ What's the best AI incantation you've seen in a real production prompt? I collect these. Bonus points if it involved offering the model money.

This article is also available in German.


AIIncentivesLearningTechnologyManagement

0 comment(s)

No comments yet. Be the first to comment.

Leave a comment

0 / 1000