📝 ESSAY

A frayed grey line severed in the middle, with a single editorial-blue strand arcing across the gap to rejoin the two broken ends, on a paper-black field

A severed line, rejoined by a single thread.

📍 IN BRIEF

How fast a team recovers from an SLA breach is decided before the breach ever happens. The organisations that recover in an hour are executing decisions they made in calm conditions. The organisations that take a week are trying to make those same decisions live, under pressure, in the worst possible conditions for making any decision at all.

Watch two organisations miss the same service level by the same margin. One has restored the service, told the customer, and closed the loop within the hour. The other is still in a meeting deciding who should send the first email. Same platform, same targets, same calibre of people, and recovery times an order of magnitude apart.

The organisations with the heaviest service level reporting are frequently the slowest to recover from a miss. If breaches were a measurement problem, more measurement would help. It does not, because all of the design effort went into the threshold, and none of it went into the moment the threshold fails.

A target without a breach path is a hope, not an agreement

Most service level design effort goes into the threshold. Almost none goes into what happens one minute after it is crossed.

Any serious approach to service level design asks a question that most organisations skip. When this commitment is breached, what actually happens next, beyond what the report will show? Who is told, who takes ownership, what does the customer hear, and what has to change before the record is allowed to close. Write those four answers down for a commitment and you have what I would call a breach path.

The cost of skipping them shows up in the first hour. Consider a trap every operations leader has watched at least once. An incident comes in as a routine priority, someone realises mid-investigation that it touches month-end payroll, and the priority is raised. The clock for the higher priority applies from the moment the incident was opened, not from the moment of the change, so the commitment can already be expired at the instant somebody upgrades it. The breach arrives by reclassification, not by neglect. A team with no pre-agreed answer spend their first 20 minutes litigating the timer, disputing when the clock really started and whether the breach really counts. A team with a breach path skip the argument entirely, because the argument was settled months ago, in a room where nobody was under pressure.

⚠️ COMMON PITFALL.

Raising an incident's priority mid-flight can put it instantly into breach, because the tighter clock applies retroactively. If your team's first response to that is a debate about the timer, the customer is watching you argue with a clock while their service is still down.

Fast recovery starts at the warning, not at the breach

By the time the timer expires, a fast-recovering team have been working the problem for half the clock.

A well-designed commitment does not stay silent until it fails. Warnings escalate through the life of the clock, so the person doing the work hears at the halfway point and again at three quarters, and the accountable manager hears at the moment of breach itself. More mature operations go a step further and forecast the misses, watching which commitments are trending towards failure across the whole portfolio rather than waiting for individual timers to expire. The percentages themselves matter less than the fact that somebody chose them deliberately, chose who hears each one, and tuned the ladder by severity. For the most severe incidents the window is so short that a half-time warning adds nothing, so the whole design concentrates on the moment of breach itself. That concentration of design effort is also a decision you can only make in advance.

The public record is unusually clear about what happens when the path was never tested. When Amazon's S3 storage service failed in early 2017, the company's own status dashboard depended on the storage that had gone down, so the page kept showing green while a large slice of the internet broke, and updates had to go out through social media instead. The alerting path had collapsed with the very service it existed to describe. Four years later Facebook dropped off the internet for hours, and its engineers found that the tools they needed to fix the network ran on the network that was down, with even physical access to facilities slowed by the same failure. The doing path depended entirely on the thing that had broken. Contrast Cloudflare's global outage in 2019, where a single change took traffic down worldwide. Service was restored in roughly half an hour because a kill switch existed and the people authorised to pull it were known, and a detailed public account followed. None of these organisations lacked talent under pressure. What separated them was the inventory of decisions available before the pressure started.

The timer is not the experience, so recover the experience

A clock that pauses generously will declare the breach over long before your customer agrees.

Service level clocks pause. They pause when a ticket sits on hold, when the work moves to a third party, when the team is waiting on the customer, and when the clock runs outside business hours. Every one of those pauses can be legitimate, and together they open a gap between the duration the timer records and the duration the customer lives through. Operations teams have a name for where that gap leads, the watermelon report, green on the outside and red in the middle. The commitment shows 99.9 per cent achievement, and the 0.1 per cent landed during the busiest trading period of the year. The dashboard celebrates while the customer quietly starts returning your rival's calls.

This is why a recovery framework cannot take what the timer says as gospel. The timer answers a single question, whether the clock ran out, while recovery has to answer three harder ones. Has the customer been told, in plain language, by you rather than by their own monitoring? Has the service genuinely returned for the people who consume it, rather than the record that tracks it? And has their confidence recovered, which you discover by asking rather than by reporting. A breach count is a lagging indicator of a relationship that was damaged weeks earlier. Chase the experience and the metric follows. Chase the metric and you will hit it while the relationship fails.

What the timer says

What the customer experiences

Clock paused while on hold with a supplier

The service has been down all afternoon

Breach avoided by the business-hours schedule

The failure ran all weekend

Resolved inside the commitment

Nobody told them it was fixed

99.9 per cent achieved this quarter

The 0.1 per cent was month-end close

A repeating breach is a design signal, not a performance problem

Recover the ticket once. If you are recovering the same commitment every month, redesign the commitment.

The first breach of a commitment is an operational event, and the breach path handles it. The fifth breach of the same commitment in a quarter is not an operational event at all, it is the design talking. A commitment that keeps failing is telling you one of three things. The target never matched the capacity underneath it, and no amount of escalation will conjure the capacity. The internal handoffs and supplier agreements stacked beneath the promise do not add up to the promise, so the top-level commitment was arithmetic fiction from the day it was signed. Or the commitment should never have existed, created for coverage rather than because anyone's decision depends on it.

The discipline that follows is the one thats often resisted, making fewer and better promises. Concentrate commitments on the touchpoints where a miss genuinely changes a customer's day, measure them to learn rather than to adjudicate, and retire the chronic breachers instead of ritually re-explaining them each month. Every hour your best people spend performing recovery on a commitment that was never achievable is an hour unavailable for the breach that is genuinely news.

Exhibit comparing two responses to the same SLA breach. The breach path, written in calm, escalates warnings at 50, 75 and 100 per cent to a named owner who tells the customer and ships a change at closure. The stopwatch, decided live, spends the breach forming a meeting, drafting the customer email and arguing about when the clock started, and closes with nothing changed.

Recovery speed is decided before the clock expires. Source: Platform Operating Manual.

Where this doesn't apply

You cannot script your way out of every failure, and a breach path is not an attempt to. It fixes the decidable in advance, the ownership, the escalation ladder, the communication obligation, the closure standard, precisely so that human judgement is free for the parts nobody could have scripted. A genuinely novel major incident will still demand improvisation, and contractual or regulated commitments carry consequences that are already written into the contract rather than yours to design. And at the other end, a commitment whose breach costs nobody anything does not deserve a breach path, it deserves deletion.

The bottom line

Take your five most important service commitments and write the breach path for each one this month, while nothing is on fire. Write down who hears the warning, who owns the response at the moment of breach as one named person rather than a rota or a committee, what the customer is told and when and by whom, and what has to change before the record is allowed to close. If a path will not fit on a page, simplify the commitment. If nobody can write it at all, be honest that what you have published is a target, not an agreement, and stop being surprised when it fails like one.

You cannot decide how to recover during a breach. You can only discover what you decided before it.