At VortexifyAI I inherited a dashboard showing a 30% improvement in a key operational metric. Leadership was pleased. I was skeptical, not of the number, but of how it was computed.
The comparison period included a system migration that changed how events were being logged, meaning two different measurement regimes had been collapsed into one trendline. The 30% was real in the data and meaningless in practice.
That experience changed how I approach any metric I didn't build myself. Before trusting a number I want to know: what changed in the data pipeline during this window, who was included in the denominator, and what would this look like split by cohort. The same instinct applies to A/B tests. A feature rolled out to high-engagement users first will always look better than it is, because you targeted people who were going to convert anyway. I built a project around correcting for exactly this: a naive comparison showed 6.45% lift, and after adjusting for selection bias the true effect was 0.9%. That's the number you make decisions with.