Winging It · Implausible Deniability
The call comes at 11:14pm on a Friday. It is always a Friday. I don't know why it's always a Friday. I have spent twenty-five years in security and I have never once been woken up on a Tuesday. Whatever force governs the timing of production incidents has a deep and abiding hatred of weekends, and I have come to accept this in the same way one accepts weather.
My phone buzzes. I look at it with the resigned calm of a man who already knows what it says. It says there is an issue. It does not say what the issue is, because the person who sent it doesn't fully know yet either. What they know is that alerts are firing, something is down, and the on-call engineer, who was halfway through something on Netflix he'll never finish, is now staring at a dashboard in his undies trying to work out whether this is a real problem or the monitoring being dramatic.
It is a real problem.
I have been this person. I have been the one staring at the dashboard. I have also been the one whose phone was on silent because I was out to dinner and had convinced myself that nothing ever happens on a Friday. Things always happen on a Friday. The human capacity for optimism on a Friday afternoon is remarkable.
Somewhere in a shared drive, in a folder that hasn't been opened since the last audit, there is a document called the Incident Management Plan. It is forty-three pages long. It has a version history. It has a RACI matrix. It has an escalation path with names, titles, and phone numbers, roughly a third of which belong to people who no longer work here. Nobody is going to open it. Nobody has ever opened it during an actual incident. It exists so that an auditor can confirm it exists, which it does, beautifully.
What actually happens is this. Someone creates a Slack channel. People join. Some of them are relevant. Some of them are not. One of them is a product manager who saw the alert and is now asking questions that will not be useful for another six hours but cannot be dissuaded from asking them now. Someone from marketing has joined by accident and is too polite to leave.
The engineer who pushed the change is contacted. This takes longer than it should because the engineer pushed the change at 4:47pm and left for the weekend with the clean conscience of someone who has never been called back. Their phone is on silent. Their laptop is closed. Friday deployments are banned in theory and happen in practice roughly every other week.
Eventually the engineer joins the call. They are calm. They are always calm. This is because engineers exist in a state of permanent, low-grade crisis and have long since lost the ability to distinguish between categories of emergency. "I pushed a config change," they say, in the tone of someone describing what they had for lunch. "It shouldn't have caused this." It did cause this. It always caused this. The config change that shouldn't have caused anything is the single most reliable source of production incidents in the history of computing.
Someone suggests rolling back. Someone else says they can't because a downstream service has already picked up the change. A third person asks what the rollback procedure is. There is one. It references a deployment pipeline that has since been replaced and a Slack channel that has since been archived. Nobody says this out loud. Everyone opens a new tab and starts Googling.
Forty minutes in, someone remembers the business continuity plan. The BCP is a magnificent piece of work. It was written over three months by a consultant named Gary who charged by the day and had a gift for formatting. It has sections on pandemic response, data centre failure, and something called "loss of key personnel," which sounds like a thriller but is about what happens when someone important goes on holiday. It was tested once, in a tabletop exercise, where everyone agreed it was comprehensive and then never thought about it again.
The BCP does not cover "an engineer pushed a config change on a Friday and now the payments service is returning errors." Nobody writes a business continuity plan for a typo in a YAML file. And yet here we are, at midnight, with twelve people on a call, because someone put a comma in the wrong place and the system has decided to express its displeasure by refusing to process anything at all.
By 1am, the problem has been identified. It was the config change. It was always the config change. The engineer fixes it. The services recover. Someone posts in the Slack channel that things are looking stable. Someone else replies with a thumbs up. The product manager asks if customers were affected. They were. The product manager asks if we know how many. We don't yet. The product manager asks when we'll know. Tomorrow. The product manager asks if there's anything they can do. There isn't. The product manager says goodnight. Everyone says goodnight. Nobody means it. The person from marketing quietly leaves the channel.
Monday brings the post-incident review. There is a facilitator. There are action items. Someone suggests banning Friday deployments. Again. Someone suggests improving the rollback procedure. Again. Someone from comms drafts a statement that describes the incident as "a brief service disruption" which is technically accurate in the same way that describing a house fire as "an unexpected change in room temperature" is technically accurate.
An action item is created to update the incident management plan. It is assigned to someone who was not on the call, does not know it has been assigned to them, and will discover it in three weeks when someone asks for a progress update. There will be no progress to update.
The incident management plan will remain unchanged. The BCP will remain untested. The rollback procedure will continue to reference a pipeline that no longer exists. And in approximately three months, on a Friday, at around 11pm, someone will push a change to production and my phone will buzz and I will look at it with the resigned calm of a man who has done this before.
I will answer it. I always answer it. And somewhere, Gary's business continuity plan will sleep soundly in its folder, undisturbed, waiting for a catastrophe that arrives as a comma in the wrong place.
— Chris
