What Google’s Gmail Outage Apology Got Right (and Missed)

Person drafting an outage message at a desk beside a clock, with an empty browser window showing a warning

A good outage apology does three things: it says what is broken, it says what you are doing about it, and it says when you will speak again. Google’s apology for the Gmail outage on 24 February 2009 did the first one well, the second one vaguely, and skipped the third.

Gmail went down for about two and a half hours that morning, for some users in the United States, Europe and Asia, according to InfoWorld’s coverage. Google’s site reliability manager wrote on the company blog that they were working hard on the problem and were “really sorry for the inconvenience.” The Gmail support site added that updates would be posted as they came in.

That is not a bad message. It is also not a complete one. And it lands at an interesting moment, because the people reading it are the same people who spent the last year learning how much they now depend on someone else’s servers.

Why do outages feel different in 2009?

A few years ago, when your email or your files went down, the server was in a cupboard down the hall and you knew the person who was swearing at it. Now the cupboard belongs to somebody else, and you find out something is wrong when a page doesn’t load.

On 20 July 2008, Amazon’s S3 storage service started returning errors at around 8:40 in the morning Pacific time. Sites that kept their images and files there, including Twitter, suddenly had broken pictures. Service in Europe was back to normal by about 3pm and in the US by about 5pm, as Network Computing reported. That is most of a working day.

Amazon later published a detailed write-up. Its explanation was that internal systems had stopped communicating properly with each other, and the fix was to take the service offline and bring it back up. Simon Willison called it honest and informative, while joking that it was the world’s longest way of saying “we turned it off and on again.”

Here’s what I take from the two events side by side. Amazon’s customers got the full account a week later. Google’s customers got a quick sorry on day one. Neither is wrong, and neither is the whole job. A good outage message needs both: the fast update while it’s happening and the proper explanation after.

What does a good outage apology need to say?

I wrote earlier that a customer’s only window into their issue is you. An outage multiplies that. Thousands of customers are staring at the same blank page, and every one of them wants the same three answers.

What is wrong. Google said Gmail was affecting “a number of users.” That’s vague, but it’s honest, because at that point nobody knew the full scope. Vague is acceptable. False is not.

What you are doing about it. “Working hard” tells people you are busy. It doesn’t tell them anything. Even “we have identified the cause and are testing a fix” gives a customer something to plan around.

When they will hear from you next. This is the one most apologies skip. “We will post updates as we have them” is a promise to talk, not a time. A time is a commitment, and a commitment you keep is worth more than any adjective in the message.

Here is the shape I’d want, written as an illustration rather than a quote from anyone:

“Mail is unavailable for some customers. We have found the likely cause and are testing a fix. Our next update will be at 11:00, whether or not we have news. We’re sorry.”

Five sentences. Every one of them gives the reader something to act on. The line about 11:00 is the one that matters, and it’s the one that takes nerve, because now you have to be there at 11:00.

Three signposts on a path, a magnifying glass, a wrench and a clock, leading a waiting crowd toward a lit lantern

Why do customers need a time and not just a sorry?

A sorry costs nothing. A timeframe has a price, because you can miss it. That is exactly why customers believe it.

When I give a customer a time, even a pessimistic one, they stop refreshing and start planning. They tell their own customers when to expect service back. Silence does the opposite. People fill it with the worst case, and now they can fill it publicly, on blogs, in forums and on Twitter.

Which means your outage is also a public event. The customer doesn’t need to call you to complain. They can tell everyone they know at once. That’s a bigger version of the old rule that an unhappy customer tells at least ten others, and it’s a reason to speak first rather than wait to be asked.

The same logic runs through internal and external SLAs. The number you promise a customer is only worth something if the team behind it can hit it. An update cadence is an SLA too. Set it, then meet it. If you promise the next update at 11:00 and the team can only produce one at noon, you promised the wrong time.

Who actually writes the message?

This is where most outages go wrong, and it has nothing to do with the apology itself. The engineers are busy fixing the problem. The support team is busy answering the phones. Nobody has been told that writing the message is their job.

So decide it ahead of time. Name one person who owns customer communication during an incident and give them two things: the authority to publish without a committee, and a direct line to whoever is fixing the problem. If the message needs three approvals, the third update will arrive after the issue is resolved and read like an obituary.

It’s the same discipline as an escalation matrix. You don’t want to work out who gets called at 2am while the phones are already ringing. You want it written down, with names, before anything breaks.

And talk to your own front line first. The agents answering calls during an outage are the ones taking the heat. Your staff don’t need sensitive technical detail to look after customers well, but they need to know what the customer-facing story is, and they need it before the customers do. An agent who learns about the outage from a caller is an agent who has just been set up to fail. It’s the same point I made about listening and communication: the quality of what reaches the customer depends on what reached your team first.

What should you do once the service is back?

Close the loop. Tell customers what failed, what you changed, and who they can reach if it happens again. The restore itself is not the end of the conversation. It’s the moment the customer decides whether to trust you with their next big day.

Amazon’s write-up is a good model for this, even with the jargon. It says what happened, in order, and what they did. You don’t need that level of detail for every incident. You do need a plain account of the cause, the fix and the change, sent to the customers who were affected, within days rather than weeks.

If the outage cost a customer real money, that conversation is also where you decide about credits or a temporary upgrade, which I’d rather offer than wait to be asked for. If you’ve got an SLA with tiered service levels, it tells you what you owe. The best operations go slightly past it, because the customer who gets more than the contract says is the one who tells people.

What does an outage teach you about your own operation?

The uncomfortable truth is that an outage is a free, involuntary audit of your communication plan. You find out in two hours what no planning meeting would show you: who writes the message, who approves it, whether the status page is somewhere customers can actually find, and whether anyone owns the next update.

After the fire is out, run a short review with a handful of plain questions:

  • How long after the first customer noticed did we say anything?
  • Did every update include a time for the next one, and did we hit it?
  • Did our front-line staff hear about it before our customers did?
  • Did we send a follow-up explaining what changed?

If you can’t answer those, you have a communication problem, and the next outage will find it for you.

Write the plan before you need it. The apology is the easy part.

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.