RE09 All articles
Engineering Culture

Goodhart's Revenge: The Metrics Your Team Follows That Are Quietly Breaking Your Velocity

RE09
Goodhart's Revenge: The Metrics Your Team Follows That Are Quietly Breaking Your Velocity

Photo: Official GDC, CC BY 2.0, via Wikimedia Commons

There's an old principle in economics called Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. It was formulated in the context of monetary policy, but it might as well have been written for engineering teams in 2024.

Here's the uncomfortable version of that idea: every metric your team is tracking is probably being gamed. Not maliciously — not even consciously, most of the time. But the moment a number becomes visible, becomes tied to performance reviews or sprint retrospectives or executive dashboards, people start optimizing for the number. And the number stops reflecting the thing you actually cared about.

This isn't a cynicism piece. It's a diagnostic one. Because the teams shipping fastest aren't the ones with the best dashboards — they're the ones who figured out which numbers to ignore.

Deployment Frequency: The Metric That Rewards Noise

Deployment frequency became a prestige metric after the DORA research got mainstream attention. High-performing teams deploy frequently. Therefore, deploying frequently makes you a high-performing team. This is a classic case of confusing correlation with causality, and a lot of engineering orgs are paying for it.

What actually happens when you tie deployment frequency to team health metrics? Engineers start breaking work into smaller chunks — not because smaller chunks are better, but because each chunk is a deploy. Features get shipped in pieces that aren't user-visible, just to move the counter. Infrastructure teams start running no-op deployments. CI pipelines get restructured around commit frequency rather than value delivery.

The number goes up. The product doesn't get meaningfully better faster.

High deployment frequency is a byproduct of a good continuous delivery culture, not the culture itself. When it becomes the goal, you get the byproduct without the culture.

Test Coverage: The Number That Measures the Wrong Thing

Test coverage is maybe the most thoroughly gamed metric in software. It's easy to measure, it sounds rigorous, and it creates a completely false sense of security.

Here's what 80% test coverage actually tells you: 80% of your lines of code are executed by at least one test. It tells you nothing about whether those tests are meaningful, whether they test real user behavior, or whether the 20% that's uncovered is the part that handles your payment processing.

Teams chasing coverage numbers write tests that pass. They don't necessarily write tests that catch bugs. You get test suites full of assertions that verify implementation details rather than behavior. You get tests written after the fact to hit a threshold. You get coverage reports that look great and test runs that provide no actual confidence.

The teams with the best actual quality signal aren't the ones with the highest coverage percentages. They're the ones asking: what would have to break for our users to notice? and writing tests around that answer.

Incident Resolution Time: The Metric That Punishes Transparency

MTTR — mean time to resolution — is a reasonable thing to care about in principle. Faster incident resolution is good. But as a tracked metric tied to team performance, it creates some genuinely counterproductive incentives.

If resolution time is measured from when an incident is declared to when it's closed, teams learn to close incidents fast. Not to resolve the underlying issue fast — to close the ticket fast. Incidents get marked resolved while the root cause is still fuzzy. Post-mortems get written to satisfy process rather than to generate insight. Engineers get cautious about declaring incidents at all, because declaring one starts the clock.

The teams that handle reliability best aren't the ones with the lowest MTTR on their dashboards. They're the ones where engineers feel safe escalating quickly, keeping incidents open until the system is genuinely understood, and writing post-mortems that are actually honest. That culture produces lower real-world resolution times — but it often looks worse on paper.

What Actually Predicts Shipping Speed

If the standard metrics are noisy, what's the signal? A few things that don't show up on most dashboards:

Time from idea to first user feedback. Not time to deploy — time to learn something real. This collapses the whole cycle: scoping, building, shipping, and getting actual signal. Teams that minimize this gap compound learning faster than teams optimizing any individual phase.

How long it takes a new engineer to ship something real. This is a proxy for system complexity, documentation quality, and cultural openness. If it takes a month for a new hire to get their first meaningful feature out, that's a structural problem that no deployment frequency metric will surface.

How often the team changes direction based on data. Velocity without direction is just motion. Teams that ship fast and adjust based on what they learn are the ones that end up somewhere worth going. This is almost never tracked.

Engineer confidence in the next deploy. Qualitative, yes. Hard to put on a dashboard, yes. But if you ask your team before a release how confident they are that it'll go smoothly and get a lot of hedging, that's a leading indicator of incidents, rollbacks, and slowdowns that your DORA metrics won't catch until after the fact.

The Audit You Actually Need

Before your next sprint planning or quarterly review, try this: list every metric your team is actively tracking and ask, for each one — what behavior does optimizing for this number actually incentivize? Not what you want it to incentivize. What it actually does.

You'll find at least two or three where the honest answer is uncomfortable. Those are the ones worth replacing — or just dropping entirely.

The best engineering teams aren't the ones with the most sophisticated dashboards. They're the ones who figured out that most dashboards are theater, and stopped performing for the wrong audience.

All Articles

Related Articles

Ghost Code: When Your Technical Debt Stops Being a Loan and Starts Being a Trap

Ghost Code: When Your Technical Debt Stops Being a Loan and Starts Being a Trap

Version Theater: The Release Ritual That's Hiding Whether Your Product Is Actually Getting Better

Ship It and Keep It: Why Your Messy First Version Might Be Your Biggest Advantage

Ship It and Keep It: Why Your Messy First Version Might Be Your Biggest Advantage