Share your voice — Be featured in Notable Men
LEADERSHIP

The Average Is Lying to You

Why your metrics might be masking the failures that matter most to your customers.

October 5, 2026 4 min read

Praveen Chebolu, Principal Software Engineer on Notable Men

Praveen Chebolu

Principal Software Engineer, Amazon

The Average Is Lying to You

The Average Is Lying to You

Early in my career, I shipped a feature that tested beautifully. Average response time was well inside our target. Every dashboard was green. We launched.

Within a week, complaints started. Not many—but the ones we got were furious. People described the product as "broken," which made no sense against our numbers. It took us an embarrassingly long time to understand what had happened: the average was fine, and roughly one interaction in twenty was terrible. Those users weren't experiencing our average. They were experiencing their own worst case, over and over, and concluding we'd shipped something unreliable.

That taught me something I've spent two decades confirming across consumer hardware, mobile platforms, and now on-device AI: The average is the most comfortable number you can look at, and comfort is exactly the problem.

Why averages hide the failures that matter

An average is a summary. Summaries work by discarding information, and the information they discard first is the tail—the small fraction of cases that sit far from the middle.

For most of what a business measures, the tail is where the damage lives.

Nobody churns because of your median experience. They churn because of the one call that went badly, the one delivery that didn't arrive, the one time the app froze during something that mattered. A customer's loyalty isn't a function of your mean performance. It's a function of your worst performance that they personally witnessed.

So when you optimize the average, you can genuinely improve your metrics while making your business worse—because the cheapest way to move a mean is usually to make good cases slightly better, not to fix the rare bad ones.

Three places I see leaders get caught by this

  • Customer experience. "Average satisfaction is 4.3 out of 5" tells you almost nothing actionable. The question is what fraction of customers had an experience bad enough to tell someone about, and what caused those specifically. Ten percent at one star is a completely different company than zero percent at one star and a slightly lower mean.
  • Hiring and performance. Teams are often evaluated on average output and managed on average competence. But teams don't fail at the average. They fail at a specific unresolved conflict, a single critical dependency on one overloaded person, or one role that's been vacant for two quarters. The averages look survivable right up until they don't.
  • Operations and delivery. "We ship on time 85% of the time" invites the wrong follow-up. The useful question is what the late 15% have in common—and whether those failures are independent or correlated. Fifteen percent scattered randomly is an annoyance. Fifteen percent clustered on your largest accounts is an existential problem wearing the same number.

That last distinction is the one I'd push hardest on. Two businesses can report identical averages and identical failure rates while having entirely different risk profiles, depending on whether their failures cluster. An average cannot tell you which one you are.

What to look at instead

You don't need statistical sophistication. You need three habits.

Ask for the distribution, not the summary — When someone brings you an average, ask what the spread looks like and what the worst tenth of cases experienced. If nobody can answer, that's your finding.

Set targets on the tail — Instead of "average response under two hours," try "95% of responses under four hours." It's a harder commitment, and it's the one your customers actually feel. Teams optimize what you measure—so measure the thing that hurts.

Go read the outliers — Not a summary of the outliers. The actual cases. Ten angry support tickets read end to end will teach you more than a quarter of aggregate sentiment scoring. The texture is the signal, and aggregation destroys texture by design.

The habit underneath

There's a broader discipline here, and it's less about metrics than about how we reach conclusions.

I once spent a week certain that a capability we needed simply didn't exist in a system because the first approach I tried was blocked. I'd found a locked door and concluded the room had no entrance. It had three. I had checked one.

Averages do the same thing to us. They hand us a single confident number, and we stop looking because the number feels like an answer. Most of the time, it's a starting point that happens to be shaped like an answer.

The leaders I've learned the most from share one trait: when a metric looks reassuring, they get more curious rather than less. They ask what the number is summarizing away. They treat a clean dashboard as a question rather than a conclusion.

That instinct is cheap to build, and it compounds. The cost of asking "what's the spread?" is about four seconds. The cost of not asking is a product your customers quietly describe as broken while every chart you own stays green.

THE SUNDAY DISPATCH

The Week's Most Considered Piece, Delivered.

One letter every Sunday. The most-read article, an editor’s footnote, and a quiet recommendation.