Your internal SLA has to be faster than the one you sell your customers, and that rule doesn’t stop at the first team you hand the ticket to. It applies to every link behind the promise: the engineers who pick up the escalation, the developers who change the code, and the vendor whose part you’re waiting on. I first wrote about internal and external SLAs here in 2008, and I said then that it sounds self-evident. It still does. And yet there’s no end of organizations I’ve dealt with where customers are offered a 4-hour SLA on a 24/7 basis while the engineers who can actually fix the problem are unavailable until the next business day. Or not even on call.
Let me state it as plainly as I did the first time. If you’re offering your customers an SLA of X hours and your engineering team (or development, or project management, or anyone else in the path) is only offering you X plus Y hours, you will lose money and you will lose customers. Not might. Will. The only open question is how many breaches it takes before it shows up in a renewal conversation. I covered the single-layer version of this in the tiered SLA matrix, where every customer cell has to sit on an internal agreement that’s faster than it. This post is about what happens after that first layer, because most of the SLA failures I’ve watched didn’t happen at the first handoff. They happened at the second or third.
Why doesn’t one fast internal SLA save you?
Say you do the hard work. Your engineering team agrees to an internal SLA of X minus Y hours, comfortably inside the customer promise. Woohoo. That will solve roughly 80% of your problems, the ones your Tier 2 group can fix on its own. The other 20% can’t be fixed by Tier 2 at all. They need a code change, a data correction or a configuration only the development team can touch, and if development is offering engineering an SLA of Z, where Z is some multiple of X plus Y, you’re still going to be in trouble on every one of them.
Here’s the uncomfortable part. That 20% isn’t a random slice of your tickets. The incidents Tier 2 couldn’t solve are, almost by definition, the hard ones, and the hard ones cluster in your Priority 1 queue at your biggest accounts. So the one link in the chain with no real commitment is the link carrying your most expensive failures. Engineering can hit its own OLA on every single ticket and your customer breaches can still climb, and this is exactly why. The engineering number is true. It just isn’t the number the customer is holding you to.
The other thing people forget is that the clock doesn’t stop at a handoff. The customer’s 4 hours started when they logged the ticket, and every queue it sits in along the way eats from those same 4 hours. “Assigned to development, awaiting triage” is still the customer’s time. That’s why the escalation matrix has to fire on the customer’s clock, not on whichever internal team happens to own the ticket right now. Measure each team on its own clock and every team can be green while the customer is red. A chain is only as fast as its slowest link.
What does the out-of-hours gap actually cost?
Do the arithmetic on the example I opened with, because it’s worse than it sounds. A 24/7 promise covers 168 hours a week. An engineering team working 9 to 5, Monday to Friday, covers 40 of them, which is 24% of the hours you’ve sold. Now log a Priority 1 at 6pm on a Friday against a 4-hour resolution SLA. If nobody who can fix it is on call, the first qualified pair of hands arrives at 9am Monday. That’s 63 hours, more than 15 times the promise, and you breached it before anyone had even looked at the ticket.
Being “on call” has to mean something, too. Someone qualified has to be reachable and able to start work inside the window, every hour you promise coverage, and that’s a staffing calculation, the same kind of arithmetic as the Erlang C staffing maths for the front line. A named engineer who might check their phone isn’t coverage. It’s hope. If you haven’t worked out who picks up at 2am on a Sunday, you haven’t got a 24/7 SLA, you’ve got a 24/7 sales brochure.

The chain doesn’t end inside your building either. If a fix depends on a hardware vendor shipping a replacement part next business day, or a software vendor with its own support tiers, that agreement is part of the chain as well. ITIL calls it an underpinning contract, and like an OLA it has to support the SLA it sits behind. A next-business-day parts contract behind a 4-hour customer promise has exactly the same problem as a development team with no SLA at all. The only difference is that now it’s someone else’s calendar you’re breaching on.
How do you build internal and external SLAs that hold together?
Work backwards from the promise, not forwards from the org chart. Take your tightest customer commitment and list every team and vendor a Priority 1 ticket can touch on its way to resolution. For each one, write down two things: its internal SLA and the hours it actually covers. Then add the chain up, including a realistic allowance for the time a ticket sits between teams. If the total doesn’t fit inside the customer promise with room to spare, you’ve found your problem before a customer found it for you.
Once you can see the gap, there are only two honest ways to close it. Either the slow link tightens, with a real commitment from the team or vendor that owns it, on-call rota included, or the customer promise loosens to something the chain can actually deliver. What you can’t do is leave it and hope. It’s imperative that your internal SLAs are better than the one you’re offering customers, and just as important that your Sales team and senior management are both on board with that. They’re the ones signing the external SLA, and they’re the ones who’ll be paying the penalties when it breaks.
That brings you to the real question: how much are you and your company willing to invest in protecting yourself from that 20%? An on-call rota for development costs money. So does a premium support tier from a vendor, and so does cross-training Tier 2 to take on fixes that currently need a developer. None of that is a support manager’s decision to make alone. It’s a business decision, and it belongs in front of senior management with the chain written out and the out-of-hours numbers attached. Show them the 63-hour Friday night ticket and I’ll show you a budget conversation that suddenly goes a lot faster.
Frequently Asked Questions
What is the difference between internal and external SLAs?
An external SLA is the service level you promise your customer. An internal SLA, usually called an OLA, is the service level one team inside your company promises another, such as engineering to support. The internal SLA has to be faster than the external one it supports, or the external one will fail.
How much faster should an internal SLA be than the customer SLA?
Fast enough that every handoff in the chain, plus the time the ticket waits between teams, fits inside the customer promise with room to spare. Against a 4-hour customer resolution, that usually means an engineering target of 2 to 3 hours, not 4. If development is also in the path, its target has to fit inside what’s left.
What happens if the development team has no SLA?
Every ticket that needs a code change has no clock on it once it leaves support and engineering. Those tickets are usually the hardest ones, so they tend to be the ones that breach the customer SLA. Get development into the chain with a written target and defined out-of-hours coverage.
Do vendor contracts need to match my customer SLA?
Yes. In ITIL terms the vendor agreement is an underpinning contract, and it has to support your SLA the same way an internal OLA does. A next-business-day vendor response sitting behind a 4-hour customer promise will breach it on every ticket that needs that vendor.
Hutch Morzaria is a CX and Support Leadership professional with 19 years of experience building and leading support organizations across SaaS, Fintech, and enterprise technology. He has held Director-level roles at Q4 Inc, AudienceView, Johnson Controls, and others, and holds ITIL Expert certification across V3 and V4.



Pingback: Why Your Internal SLA Must Outpace Your Customer SLA: A Guide to Avoiding Costly Mistakes - CX Master