A tiered SLA matrix is the honest answer to a problem every support leader runs into sooner or later: you can’t give every customer platinum service, so you decide in advance who gets what, for which kind of problem. Mine puts three customer tiers across the top, three incident priorities down the side, and a response time and a resolution time in each of the nine cells. I first published it here in 2007. The numbers are below, and they still hold up as a starting point. What’s changed since then is the tooling, not the logic. Every ticketing system worth paying for can now start two clocks on a ticket the moment it’s logged, and all of them still need somebody to decide what those clocks should say. That decision is the matrix.
You obviously want to offer every customer the premier, best in the world, platinum level of service. Unfortunately, that doesn’t always make financial sense. Customers need to be tiered by the amount of money they pay you, which is the whole argument behind the 80/20/30 rule, and incidents need to be tiered by the impact on their business. Put those two axes together and you get a grid instead of a single promise. That’s the part most contracts skip. A blanket “we respond within one hour” treats a total outage at your biggest account exactly the same as a cosmetic bug at your smallest, and one of those two customers is going to be badly served. Usually both are. The grid fixes that by making the trade-off on paper, before the phone rings, instead of leaving it to whoever happens to pick up the ticket.
The tiered SLA matrix, cell by cell
Here’s the matrix. It’s a basic version, and you’ll need to customize it for your business rather than copy it cell for cell.
| Tier 1 | Tier 2 | Tier 3 | |
|---|---|---|---|
| Priority 1 70%+ service impact or total loss of service |
5 min response 4 hr resolution |
30 min response 8 hr resolution |
1 hr response 24 hr resolution |
| Priority 2 50% to 70% service impact |
15 min response 8 hr resolution |
1 hr response 24 hr resolution |
4 hr response 72 hr resolution |
| Priority 3 Up to 50% service impact |
30 min response 12 hr resolution |
4 hr response 72 hr resolution |
24 hr response 96 hr resolution |
Read it from top left to bottom right. A Tier 1 customer with a Priority 1 incident sits in the top-left corner: 5 minutes to respond, 4 hours to resolve. A Tier 3 customer with a Priority 3 incident sits bottom right, with 24 hours to respond and 96 hours to resolve. Every other cell lands somewhere in between, and each step away from that top-left corner buys your team more time. The priority rows are defined by how much of the customer’s service is affected. I like percentages here because they take the argument out of classification. “It’s really urgent” is an opinion. “Two-thirds of our users can’t log in” is a number, and the number decides the priority before anyone has to negotiate it. It also hands your agents a question to ask on the first call, rather than a judgement to make on their own while an angry customer waits on the line.
What’s the difference between response time and resolution time?
Response is how long it takes someone qualified to pick the incident up and start working it. Resolution is how long it takes to fix it. Look at the gap in the matrix: a Tier 1 Priority 1 incident gets a 5-minute response but a 4-hour resolution, which is 48 times longer. That gap is deliberate. Problems generally can’t be resolved on first contact, because the work to analyze the problem takes time. Don’t kid yourself otherwise.
The mistake I see is a team that promises a fast response and then gets judged as if it had promised a fast fix. The customer hears “5 minutes” and expects the problem gone in 5 minutes. So the contract, and every status update you send, has to be explicit about which clock is running. Use that difference to your advantage. A quick, competent response buys you the time the resolution genuinely needs, as long as the customer can see work happening. Silence is what turns a 4-hour fix into an escalation.
The response column is also a staffing decision in disguise. A 5-minute response for your top tier means someone qualified is free to pick up within 5 minutes, every hour you promise coverage. That isn’t a policy question. It’s arithmetic, and it’s the same Erlang C scheduling maths that turns call volume into agents per hour. Run your response targets through it before you sign them. If the headcount it spits out isn’t in the budget, the target has to change. Wishing won’t move a queue.
Your SLA can’t be tighter than your OLA
This is the point from the original 2007 version I’d underline twice. Your customer-facing SLA has to be achievable, and it’s only achievable if the internal agreements sitting behind it are faster. An OLA, or operational level agreement, is the service level your own internal departments offer each other: how quickly engineering picks up an escalation, how quickly the network team restores a link, how quickly a vendor ships a patch. HDI describes the OLA as the internal contract that sets what each group has to commit to so the SLA can actually be delivered. That’s the right way to think about it.
Walk it through the matrix. You promise a Tier 1 customer a 4-hour resolution on a Priority 1 incident. Your engineering team’s internal target for a production defect is 8 hours. You’ve just signed a promise that nobody inside your company has agreed to keep. A customer-facing SLA that’s more stringent than your own OLA is doomed to failure, and depending on how the contract is written, some fairly large financial repercussions come with it. Service credits get expensive fast on your biggest accounts. They’re also the accounts sitting in the tightest cells.
So build it in the right order. List your OLAs first, per priority, and agree them in writing with the teams that own them. Then set each cell of the customer matrix with enough room left over for the handoffs between those teams. If a cell can’t be supported, either the OLA tightens with a real commitment from the team that owns it, or the customer promise loosens. Don’t publish a cell with nothing underneath it. The matrix is only as honest as the slowest team it depends on.

Frequently Asked Questions
What is a tiered SLA?
A tiered SLA sets different response and resolution targets depending on who the customer is and how badly an incident affects them. The matrix above uses three customer tiers based on revenue and three priority levels based on the percentage of service impacted, which gives nine separate targets instead of one blanket promise.
What is a good P1 response time?
It depends on the customer tier and on what you can staff. In this matrix a Priority 1 incident gets a 5-minute response for Tier 1, 30 minutes for Tier 2 and 1 hour for Tier 3. The right number is the fastest one your team can hit every hour you promise coverage.
What’s the difference between an SLA and an OLA?
An SLA is the service level you promise your customer. An OLA is the service level your internal teams promise each other so that the SLA can be met. The OLA has to be tighter than the SLA it supports, or the SLA will fail.
How do you prioritize incidents for an SLA?
Tie priority to measurable business impact, not to how urgent the customer says it is. Here, 70% or more of service impacted, or a total outage, is Priority 1; 50% to 70% is Priority 2; anything up to 50% is Priority 3.
Hutch Morzaria is a CX and Support Leadership professional with 19 years of experience building and leading support organizations across SaaS, Fintech, and enterprise technology. He has held Director-level roles at Q4 Inc, AudienceView, Johnson Controls, and others, and holds ITIL Expert certification across V3 and V4.


