The Password That Burned
(The Ticket That Still Haunts Me after 15 years)
Early in my 1st-level support days I got the classic "we lost the admin password" ticket.
Business-critical application. Low-level software. No backdoor, no recovery flag, no magic command. The only way forward was the one we always gave: reinstall and restore the data from backup. I sent the standard template, closed the ticket in my head, and moved on.
This customer didn't close it.
He pushed back hard, so I scheduled a call. On the line he sounded desperate in a way most tickets never reach. I walked him through the security reasons again, step by step, expecting the usual frustrated "fine, we'll do the reinstall." Instead, he went quiet for a long second and then told me the real story.
One of his colleagues - the only guy who managed the entire infrastructure - had died in a car accident a few weeks earlier. The laptop with every password, every recovery key, and every piece of documentation went up in flames with the car. No off-site copy. No shared vault. No second person who knew the master credentials. The whole company's core systems were now running in a zombie state: accessible to nobody, restorable by nobody.
I didn't know what to say. There was nothing technical left to offer. The ticket wasn't about a forgotten password anymore. It was about a single point of failure that real life had just executed with zero mercy.
That one call stayed with me longer than any outage I've ever worked.
Because the lesson isn't just "back up your passwords." It's bigger and darker and hits every layer of support, engineering, and leadership.
Never let the most critical knowledge live in one person's head or one device's hard drive.
In support we see this pattern constantly. The senior admin who "just knows" how the legacy system works. The architect who keeps the only copy of the disaster-recovery runbook on their laptop. The one engineer who can still talk to the ancient mainframe because everyone else who understood it retired or moved on. We call them heroes until the day they get hit by a bus, quit, or simply go on vacation and the pager starts screaming.
The fix is brutally simple and still wildly unpopular in companies that reward individual brilliance over durable systems:
π Force shared credential vaults and rotate access regularly.
π Document the ugly stuff even when it feels boring.
π Build runbooks that any competent person can follow, not just the one genius who wrote them.
π Make knowledge transfer part of every role's actual job, not an optional nice-to-have before someone leaves.
π Test the process. Pretend your best person is gone tomorrow and see if the company can still operate. Most can't. That's the real red flag.
The tragic part? The customer in that story didn't fail because of bad software. They failed because they treated one human being as their entire business continuity plan. And real life has no rollback.
Support teams see these single points of failure every single day. The best ones don't just fix the immediate ticket - they quietly push the company to fix the underlying fragility before the next car accident, layoff, or surprise resignation turns a manageable problem into a company-killing one.
Leadership that listens ends up with systems that survive people leaving.
Leadership that doesn't ends up with outages that start with "we lost the password" and end with lawyers.
We are all one unexpected event away from becoming that customer on the phone.
Build like you know it.
Ever seen a single point of failure finally bite a company - or watched a team finally fix theirs before it did? Drop the real (and painful) stories below π - the more honest, the more useful for everyone else still carrying those keys in their head.
This article is also available in German.
TrenchOps π
0 comment(s)
No comments yet. Be the first to comment.
Leave a comment