At half past one in the morning, Hannah’s phone lit up on the bedside table with a notification from her bank.
£18.50, paid to her own company.
That part was normal. Hannah had been customer number seven of the coffee subscription business she now ran the technology for. She’d signed up on launch day to test the checkout, and she’d never cancelled, partly out of loyalty and partly because the coffee was good. Every month, in the small hours, the renewal job billed her along with everyone else. She usually slept through it.
She was awake that night only because it was the night the clocks went back, and she’d stayed up reading, enjoying the extra hour the way you do.
An hour later, just after the clocks rolled back, the phone lit up again.
£18.50, paid to her own company.
She picked it up and looked at the two notifications, one above the other. Same amount. Same merchant. And the same timestamp on both of them: 01:30.
It took her a few seconds to understand what she was looking at. When she did, she sat up very straight in bed, because she knew that if it had happened to customer number seven, it was happening, at that moment, to everyone.
Hannah isn’t one person. Like Daniel and Priya from the last few weeks, she’s a composite of the CTOs and heads of technology I’ve sat across from over the years, built from real conversations so the story holds together. But every beat of what happens to her next has happened, in one shape or another, to someone real. Some of it more than once.
Here’s the business, so you can feel the size of what was about to go wrong.
A specialty coffee subscription, based in Bristol, about twelve thousand four hundred active subscribers. Nothing enormous, but a real business with real margins, a warehouse team that started picking orders at six every Monday morning, and a board that had just approved a second roastery. Hannah ran technology with exactly one in-house engineer, Josh, who was very good and very tired, and who was that week on his first proper holiday in two years, in Lisbon.
The platform itself had been built by an agency about two years earlier. Good work, by everyone’s account. When the build finished, the agency did what agencies do: they moved the account across to their support desk and moved their best people on to the next project. The lead developer, a man called Marcus who understood every corner of the system, had left the agency altogether the following spring.
There was a support contract. Hannah had read it when it was signed. It promised a one-hour response, around the clock, for any P1 incident. It defined a P1, in the kind of clause nobody reads twice, as “complete or substantial unavailability of the service.”
There was monitoring, too. A tidy dashboard with green lights for the website, the checkout, the database, and the payment gateway. Hannah opened it on her phone at 01:33, with the two notifications still on her screen.
Every light was green.
And they were right to be. The site was up. The checkout was working. The database was healthy. The payment gateway was processing payments quickly and without a single error. By every measure anyone had thought to put on that dashboard, the system was working perfectly.
It was working perfectly at charging twelve thousand people twice.
It took the team most of the following week to reconstruct it, so let me give you the short version now.
The renewal job was set to run at 01:30 every night. Whoever set it up, two years earlier, had set that time in UK local time rather than in UTC, which is the kind of small, reasonable-looking decision that nobody would ever flag in a code review. On the night the clocks go back, the hour between one and two in the morning happens twice. The scheduler saw 01:30 arrive, ran the job, and then an hour later, when the clocks rolled back and 01:30 arrived again, it did exactly what it had been told to do and ran it again.
The renewal job took about two hours to work through every subscriber, oldest accounts first. Hannah, as customer number seven, had simply been one of the first people billed in the second run.
Which meant that as she sat there in bed, the job was working its way down the list at roughly a hundred customers a minute, and would keep going until it reached the end at around half past three.
Every minute she spent working out who to call was another hundred customers charged twice.
She rang the support line first, because that’s what the support line was for.
A polite, calm voice picked up within a few rings. She explained. He listened. Then he asked the question that was, in his defence, the question his job required him to ask.
“Is the site down?”
It wasn’t.
“Can customers log in and place orders?”
They could.
He was sorry. He could hear this was serious. But under the contract it wasn’t a P1, because the service was available. He’d log it as a P2, which meant it would be picked up by the platform team at the start of the next business day. Hannah asked when that was. It was Sunday morning. The next business day was Monday at nine, a little over thirty-one hours away.
She asked if he could stop the job himself. He couldn’t. He didn’t have access to the scheduler for her account. Only the platform team did, and the platform team worked business hours.
I want to be fair to that man, because it would be easy to make him the villain of this story, and he isn’t. He did his job exactly as it had been written. The problem was that his job had been written around a single question, is it down, and the thing happening to Hannah’s business that night wasn’t down. It was wrong. Nobody had ever written a contract, a dashboard, or a job description for wrong.
She called Josh. It rang out. She sent him a message, then another. Nothing. It was nearly two in the morning in Lisbon too, and he was asleep, as he had every right to be.
She searched her inbox for Marcus and found an address at the agency that bounced. She found him on LinkedIn and sent a message she knew he wouldn’t see until morning, if ever.
Then she found the handover document. Fourteen pages, written the week the project closed. Most of it was screenshots of the admin panel. Hannah scrolled through it on her phone, looking for anything about scheduled jobs, and found exactly one line: “Background jobs run on the worker server.”
That was it. Nothing about where the worker server lived. Nothing about how to stop a job once it started. Nothing about what to do if it did something it shouldn’t. The document had been written by people who knew the system so well that the important parts had never occurred to them as things that needed writing down.
It was 02:04. Somewhere around three and a half thousand customers had now been billed twice, and the number was still climbing.
Josh called back at twenty past two. His phone had been face down on a hotel bedside table, and the fourth vibration had finally woken him.
He understood inside about fifteen seconds. He also had a problem of his own: he’d left his laptop at home, deliberately, because it was his first holiday in two years and he’d promised his partner. All he had was his phone, and the hosting console wanted a hardware security key he didn’t have with him.
But Hannah had the admin credentials, sitting in the company password manager, for a console she had never once logged into.
So Josh talked her through it. Which menu. Which project. Which of the four servers with nearly identical names was the worker. At one point she found a large button marked “Restart” and asked if she should press it, and she heard him sit up in bed a thousand miles away. Don’t touch that, he said. A restart re-queues pending jobs. You’d make it worse.
It took them twenty-seven minutes, on a crackly hotel line, a head of technology who’d never touched the infrastructure being guided by an engineer who couldn’t, to find the right setting and scale the worker down to nothing.
At 02:47 the job stopped.
Hannah checked her bank app. Still two notifications. No third. She sat on the edge of the bed for a while and didn’t sleep again that night.
By Monday lunchtime, they had the numbers.
Just under eight thousand customers had been charged twice, about £147,000 taken that shouldn’t have been. The money itself could be refunded, and was, though refunding eight thousand card payments is its own small project with its own fees. The support inbox received over four hundred emails before ten on Monday morning, some confused, some furious, a few from people who’d gone overdrawn and been hit with charges by their own banks. Some customers didn’t email at all. They went straight to their bank and raised a chargeback, which costs a merchant a fee every time whether or not the merchant was at fault, and counts against them with their payment provider. And over the following fortnight, a little over two hundred customers quietly cancelled.
None of that is catastrophic for a business that size. But it was expensive, it was embarrassing, and it landed a fortnight before a board meeting about a second roastery.
The chair asked the question Hannah had known was coming since she sat up in bed on Sunday morning.
“Who owned this?”
Hannah told me later that she’d spent most of Monday night trying to answer that question fairly, and the honest answer she kept arriving at was: nobody.
Not nobody in the sense that people had been careless. Nobody in the sense that every single party had done precisely what they’d been asked to do, and the failure lived entirely in the space between them.
The agency had built the system well, handed it over, and moved on, which is what they were paid for. The support desk had applied its contract exactly as written, and by the contract’s own definition, nothing was wrong. The monitoring had reported, accurately, that everything it had been built to watch was healthy. Josh had been on a holiday he’d booked months in advance and was entitled to. Marcus had left for another job, as people do. And the one decision that actually caused the incident, a scheduled time set in local time rather than UTC, had been made two years earlier by someone nobody in the room could name, for a reason nobody could remember, and had then sat there quietly doing nothing wrong for two years, until the one night it did.
There was a moment in that meeting where one of the non-executives suggested, fairly gently, that the support vendor should be replaced and that perhaps someone ought to take responsibility.
This is the part of the story where Hannah wins, and I want to tell it carefully, because it would have been so easy for her to take the other road.
She could have fired the vendor. It would have been defensible, it would have felt like action, and the board would have been satisfied. Instead she said something close to this: if we replace the vendor and change nothing else, we’ll be back in this room in eighteen months, having the same conversation about a different Sunday. The vendor didn’t fail. Our ownership did. Give me sixty days to fix that, and I’ll come back and show you it’s fixed.
They gave her the sixty days.
I want to set this out properly, because it’s the useful part, and because none of it required a bigger budget. It required a few uncomfortable questions and the patience to act on the answers.
First, she found out what she actually had. Before changing anything, she ran a test. She asked two people, separately, and without warning, the same question: if our live product started doing something wrong, not down but wrong, at two in the morning on a Sunday, whose phone rings, and what’s the first thing they do?
She asked the vendor’s account manager. The answer was: you’d call us, and we’d log a ticket.
She asked Josh. His answer was: out of hours, that’s the vendor.
Each of them named the other. That was the whole diagnosis in two sentences. Both answers were reasonable, both people were competent, and between them they described a system in which nobody’s phone rang at all.
Second, she changed what counted as an emergency. The old question was, is it down? The new one was, is it taking money, sending things, or telling customers things it shouldn’t? If the answer to that was yes, it was the top priority, whatever the green lights said. Hannah put it to her board in one line I’ve since borrowed more than once: when the site goes down, we lose some sales for an hour. When the site goes wrong, we lose customers’ trust and create liabilities, and we might not notice for a very long time. Down is loud. Wrong is quiet. The contract had been written entirely around loud.
Third, she put a name on it. Not a rota, not a shared inbox, not a support tier. One named person, accountable for the live product, who knew why the system had been built the way it had and who would still be there next year. She moved the running of the platform to a provider that worked that way, and the first thing she asked their named lead to do was read every line of code that moved money or talked to customers, and tell her what could go wrong at two in the morning.
Fourth, she made stopping things boring. Every job that took money, sent emails, or touched orders now had a written, tested way to stop it, in plain language, on a single page, that someone who’d never seen the infrastructure could follow at two in the morning on a phone. And she tested it on exactly that person. She gave the page to her finance director, who has never written a line of code, set up a harmless test job, started a timer, and asked him to stop it. It took him six minutes. She told me that was the moment she started to sleep properly again.
Fifth, she went looking for the next Sunday before it arrived. The new named lead audited every scheduled job in the system. They found eleven set in local time. Eleven small, reasonable-looking decisions, each one waiting for a particular night of the year. All eleven were moved to UTC within the month.
When the sixty days were up, Hannah went back to the board. She didn’t bring a slide about the vendor. She brought the answers to the two-person test, asked again, and this time both people gave the same name.
The clocks went back again the following October.
Hannah was awake again. She admits it was on purpose. She lay there with her phone on the bedside table and watched it, the way you might watch a pot you’re fairly sure won’t boil over.
At 01:30, it lit up. £18.50, paid to her own company.
She waited. The clocks rolled back. 01:30 came around a second time.
Nothing.
Two minutes later, one more message arrived, this time from the named lead at the new provider. It was just after seven in the morning where he was, in Kolkata, and he’d been at his desk for a while.
Morning from our side. Clocks went back on yours. All eleven jobs ran once, as expected. Nothing for you to do. Go back to sleep.
She did.
I love that ending for a reason that has nothing to do with Kolkata being where my own company happens to be. It’s that the most important thing that happened that night was that nothing happened, and that someone who owned it was awake to confirm it. That’s what good ownership looks like from the outside. It’s almost completely boring. You only ever see it in the absence of a story.
I’ve been on the other end of Hannah’s phone call.
Years ago, on a Friday night, somewhere around two in the morning, my phone rang. It was a client of ours, John, calling from Dallas, and he was not happy. His database had gone down in the middle of a demo to his investors. He told me he’d been left with egg on his face in front of the people funding his company.
I called our delivery head. Within ten minutes he called me back: he, the project manager and the team would be in the office by six thirty. They worked from seven until two in the afternoon, and John woke up to a working system.
John stayed with us for a long time after that, and eventually became about thirty percent of our monthly revenue. But I’ve never believed he stayed because the database got fixed. Databases get fixed. He stayed because when it broke, one person picked up, and one person was accountable at six thirty in the morning, and he never once had to explain his problem to someone meeting it for the first time.
That’s the whole reason we run managed services the way we do. One named person, from the start, who stays for the life of the engagement and is personally accountable for whether the thing keeps working. Not because it sounds good in a proposal, but because I’ve seen, from both ends of the phone, what happens in the space where nobody owns it.
I’d be doing exactly the thing I’m warning you about if I told you every business needs a managed services provider. Plenty don’t.
If you have an in-house team with a real on-call rota, a runbook that a non-engineer can follow, and two people who’d give you the same name at two in the morning, you already have what Hannah built in her sixty days. Keep it, look after it, and don’t let anyone sell you a replacement for something that works.
If you’re earlier than that, still changing the product every week, you probably want a dedicated team that builds and runs together, not a support arrangement at all. And if you’re live, stable, and simply don’t have the people to watch it, that’s where an arrangement with a named owner and a proper service level earns its money. The question that decides which of these you need isn’t in our rate card. It’s this one: what would one hour of your system being wrong cost you? Not down. Wrong. Most people have never priced it. Hannah priced it at roughly £147,000 and two hundred customers, after the fact.
So here’s what I’d ask you to do this week, and it takes about ten minutes.
Pick two people. One on your side, and one on your vendor’s side if you have one, or two people on your own team if you don’t. Ask each of them, separately and without warning, the question Hannah asked: if our live product started doing something wrong, not down but wrong, at two in the morning on a Sunday, whose phone rings, and what’s the first thing they’d do?
Then compare the answers.
If you get the same name and the same first step, you have an owner. Well done, and you can stop reading here.
If you get two different names, you have what Hannah had: each side believes the other one has it covered. If you get a job title or a team name instead of a person, “support handles it,” then nobody is holding it at two in the morning. A title doesn’t wake up. And if you get a pause, followed by “that’s a good question,” then you already know.
And one more thing, because it’s a gift of timing. In the UK, the clocks go back on Sunday 25 October. In the US, they go back a week later, on Sunday 1 November. Between now and then, ask whoever looks after your systems a single question: is anything that moves money, sends messages, or touches orders scheduled in local time? It’s a short question with a short answer, and it’s a great deal cheaper to ask in September than at half past one on a Sunday morning.
If you run the test and you don’t like the answers, send them to me. I’ll tell you plainly whether you have an ownership gap, and what it would take to close it, even if the honest answer is that your own team can close it without us.
Hannah’s phone stayed dark that second October. I’d like yours to as well.