The Incident Manager’s Real Job Is Communication

Incident manager at a central podium passing a glowing line of communication between engineers and waiting customers

The incident manager’s real job is communication. Fixing the fault belongs to the engineers. The incident manager owns everything around it: who knows what, who decides, and when the customer hears from you next. If that sounds like a soft skill, it isn’t. A customer can forgive a fault they understand. It is much harder to forgive one nobody will explain.

I covered the customer-facing half of this in what Google’s Gmail apology got right and missed: say what is broken, say what you are doing, and say when you will speak again. This post is about the other half. Who makes that happen, and where does the role sit in ITIL?

What is an incident, and what does ITIL say about managing one?

ITIL defines an incident as an unplanned interruption to an IT service, or a reduction in its quality. The goal of incident management is to restore normal service as quickly as possible and keep the impact on the business as small as possible. In ITIL version 3, published in 2007, incident management sits in the Service Operation stage of the lifecycle, alongside problem management, event management and request fulfilment. The IT Service Management Forum is the community most practitioners go to for the detail.

The point most teams miss is in that goal. “Restore service” is the technical target. “Minimize the impact on the business” is a communication target as much as an engineering one. A customer who knows the fix is 40 minutes away can reschedule a meeting. A customer who knows nothing cannot.

Incident management also stays deliberately narrow. It restores service, even if that means a workaround. It does not find the root cause. That is problem management’s job, and mixing the two is how an outage that should take an hour turns into a day-long debate about why it happened while the service is still down.

What does an incident manager actually do?

In ITIL terms, the incident manager owns the incident management process. Day to day, for a serious incident, that means taking charge of the response without taking over the fix. Four jobs, in order:

  • Declare it and set the priority. ITIL derives priority from impact and urgency: how many users or how much of the business is affected, and how quickly it needs fixing. The incident manager makes the call that something is a major incident, because that call triggers a different, faster procedure.
  • Get the right people on it. That is functional escalation, pulling in the technical group that can fix it. It is the same logic as an escalation matrix, which is worth having written down before 2am rather than after.
  • Escalate upward when needed. Hierarchical escalation is telling management early, not late, when an incident is going to breach what was promised.
  • Own the communication. Agree the facts, write the update, publish it on the time you promised, and make sure the people answering the phones have it first.

The last one is the job that gets dropped. The engineers are busy fixing, and the support team is busy answering. Unless the incident manager owns the message, nobody does.

Why does communication belong to the incident manager and not the engineers?

Because the two jobs pull in opposite directions. An engineer deep in a fault wants silence and focus. A customer wants a steady stream of updates. If the same person has to do both, the updates are the thing that slips, and they slip first when the problem is worst, which is exactly when customers need them most.

The split is simple. The engineers talk to the incident manager. The incident manager talks to everyone else. The service desk, which ITIL treats as the single point of contact for users, relays what the incident manager has agreed. Nobody on the fix team gets interrupted for a status request, and nobody outside it has to guess.

The same idea runs through a customer-facing SLA. The internal and external SLA targets you set only matter if someone is actually watching the clock against them during the incident. That someone is the incident manager.

What happened in a real incident?

On 1 September 2009, Gmail was down for about 100 minutes. Google explained afterwards that some servers had been taken offline for routine upgrades, and that it had underestimated the load some recent changes put on its request routers. Those routers became overloaded in turn, and within minutes nearly all of them were affected, so users couldn’t reach their mail through the web interface. IMAP and POP access kept working because they don’t use the same routers. InformationWeek’s report has Google’s engineering VP calling the outage “a Big Deal,” and a full explanation was published the same day.

Compare that with the February outage I wrote about. In February, Google gave a quick apology and promised updates. In September, it gave a quick apology and a detailed cause within hours. Whether or not anyone at Google carries the title of incident manager, that is incident-manager behavior: fast acknowledgement, a named owner of the explanation, and a written account that treats customers as adults.

The 100 minutes of downtime is a failure of change management and capacity planning. What the customer thinks of Google a week later is a failure or success of incident communication, and that part is entirely within the incident manager’s control.

Four stages of incident handling on stepping stones around a lantern: alert, handover, fix and review

How does this connect to problem management and the SLA?

An incident ends when service is restored. The conversation shouldn’t. Three things should follow:

  1. A plain account to affected customers: what failed, what you changed, who to contact next time. If you’ve done the tiered service levels properly, it also tells you what you owe the customers who were affected.
  2. A handover to problem management for root cause. The incident manager’s record of the timeline is the best input that team will get.
  3. A short review of the communication itself: how long before the first update, did every update include a next-update time, did the front line hear before the customers did?

That third step is the one that separates an operation that learns from one that keeps repeating the same bad night.

What should you do this month?

Three things, none of them expensive. Name your incident managers in advance, including who covers nights and weekends. Write a one-page major incident procedure with the update cadence and the template wording in it. Run one tabletop exercise where the only thing you practice is the communication.

If you only do one, name the person. Most of the other failures follow from there being nobody whose job it was. As I said in the piece on exceptional customer service, the customer’s only window into their issue is you. The incident manager is the one holding that window.

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.