The Importance of Measurements: How to Actually Track Your KPIs (Part 2)

Kpi4

The importance of measurements isn’t really about the formula. It’s about capturing the same data, the same way, every single time, so the number from March still means something in October. My previous post on choosing the right KPI covered why you pick one from the customer’s side of the table, not yours. This one is about how to measure KPIs once you’ve made that choice: what to capture, where a spreadsheet stops being enough, and why the average sitting on your dashboard right now is probably lying to you a little.

Start with whatever you’ve decided actually matters. Outage minutes. First call resolution. Ticket volume against a specific date. Get it into a spreadsheet or whatever tracking tool you already have, then keep adding to it every time the thing happens again. Consistency is the whole job here, more than the tool is. A number measured one way in January and a slightly different way in June isn’t a trend, it’s noise wearing a trend’s clothes, and no amount of formatting fixes that after the fact.

Monthly availability chart broken into Availability, Not as Bad, and Bad Stuff

How Do You Actually Measure KPIs?

Capture more than the headline figure. If you’re tracking outages, don’t just log that one happened. Log who it hit, how long it ran, how many calls or tickets it generated, and whether the cause was internal or a third party. That last tag sounds small and it isn’t: a pattern of third-party outages means an uncomfortable vendor conversation, and a pattern of internal ones means an uncomfortable one with your own team. Either way, you need the tag before you can have either conversation, and you can’t go back and reconstruct it six months later from memory.

A spreadsheet is genuinely fine for this in the early stages. Most teams reach for a fancier tool before they’ve proven the discipline of capturing clean data consistently, which is backwards. Get the habit right first. Once you’re logging the same fields, the same way, for every single incident, you’re ready to think about whether the tool itself needs to change. Not before. A polished dashboard built on top of inconsistent data is just a better-looking version of the same problem.

This matters more than it sounds like it should, because the team that skips this step almost always finds out the hard way, usually in a leadership meeting where two people are looking at the same metric and getting two different numbers out of it.

First call resolution is a good example of how this goes wrong. Two teams can both report a 78% FCR and mean completely different things by it. One counts only calls with no follow-up contact within 24 hours. The other counts anything closed on the first interaction, regardless of what happens next. Neither definition is wrong on its own. What’s wrong is comparing the two numbers as if they measure the same thing, which is exactly what happens the moment they land on the same slide.

Why a Spreadsheet Stops Working

It stops working the moment incident volume outgrows one person’s ability to keep it updated by hand. At a 210-person support organization, that point arrives faster than most people expect. A handful of outages a month is a spreadsheet problem. A few a week is an incident-and-problem-management problem, and pretending otherwise just means someone is manually reconciling three versions of the same tracker every Friday afternoon instead of doing anything useful with the data.

This is where a proper incident and problem management system earns its keep, tied to the same ITIL discipline that governs how incidents get classified, escalated and closed. I saw this play out directly during the AudienceView merger integration: once P2 tickets started getting tracked with the same rigor as P1s, same fields, same cause tagging, same review cadence, average resolution time on those tickets dropped from over 20 days to under 2. That wasn’t a tooling win. It was a consistency win that a better tool happened to make possible, and the tool would have changed nothing on its own.

If Six Sigma is part of your world, the DMAIC structure, define, measure, analyze, improve, control, is the same discipline in different language. I cover the basics of Six Sigma properly elsewhere, but the short version holds here too: you can’t improve what you haven’t measured the same way twice, and a tool can’t fix a measurement process nobody agreed on first.

The review cadence matters as much as the capture. A tracker nobody looks at until the quarterly business review is just data that happened to get saved. Pull it into a short weekly look, even ten minutes, and patterns show up fast: the same vendor causing three of your last five outages, or first call resolution quietly sliding while nobody was watching the daily number closely enough to notice.

Why Averages Lie

There’s an old line, usually pinned on Mark Twain, sometimes on Disraeli, nobody’s actually sure, about three kinds of lies: lies, damned lies, and statistics. Most people already carry a low-grade distrust of any number presented as fact, and averages are where that distrust is most deserved. An average is the easiest number in the world to present confidently and the easiest one to be wrong about.

An average on its own tells you almost nothing. One outage that runs four hours will drag your uptime number down further than ten five-minute blips combined, and if you’re only watching the monthly average you’ll miss that your real problem is one bad vendor, not a systemic one. Good analysis means setting a target, then actually checking whether the changes you make move that target, not just watching the top-line number and assuming it tells the whole story on its own.

This is exactly the trap dashboards fall into at the executive level too. A metric that’s technically correct but defined inconsistently between teams looks fine right up until two departments report different numbers for the same thing in the same meeting, and somebody has to explain why. I’ve written about that failure mode from the BI side here, if you want the fuller version of how it happens further up the org chart.

A better habit than trusting the average outright is to also look at where the 90th percentile sits. If your average resolution time is two hours but your 90th percentile is eighteen, most of your tickets are fine and a smaller group is badly broken, and that’s a completely different problem to solve than a uniformly slow process would be. The average would never have told you which one you actually have.

An average is a summary, not a verdict. Somebody still has to know what’s underneath the number before they act on it.

Frequently Asked Questions

What’s the first step in measuring a KPI?
Decide what you’re actually tracking and capture it the same way every single time, even if it’s just a spreadsheet to start. Consistency matters more than the tool in the first few months.

Why do average KPI numbers hide real problems?
A single extreme outlier, like one long outage, can pull a monthly average down further than several small ones combined. Looking only at the average hides which type of incident is actually driving the trend.

When should a support team move from a spreadsheet to an incident management system?
Once outage or ticket volume outgrows what one person can keep updated accurately by hand, usually somewhere around a handful of incidents a week rather than a month. A 210-person support organization typically hits that point well before smaller teams do.

What’s the difference between measuring KPIs and improving them?
Measuring tells you what’s happening. Improving requires a defined process like DMAIC, define, measure, analyze, improve, control, to act on what the data shows. You can’t do the second reliably without the first being consistent.

1 thought on “The Importance of Measurements: How to Actually Track Your KPIs (Part 2)”

  1. Pingback: KPI's and the Importance of Measurements (part 2) - CX Expert

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.