How I approach every product question, before touching data.
1
Define the metric
Start by asking what success looks like and what it should not hurt. Primary metric, guardrail metrics, before any code.
2
Design the test
Structure for causal clarity. Control for novelty effects, check for interference, size for the lift that actually matters.
3
Decide with confidence
Translate p-values into plain business language. Tie statistical significance to magnitude, risk, and rollout cost.
The number that looked great and wasn't
On naive lift, selection bias, and why I stopped trusting simple comparisons
▾

At VortexifyAI I inherited a dashboard showing a 30% improvement in a key operational metric. Leadership was pleased. I was skeptical, not of the number, but of how it was computed.

The comparison period included a system migration that changed how events were being logged, meaning two different measurement regimes had been collapsed into one trendline. The 30% was real in the data and meaningless in practice.

That experience changed how I approach any metric I didn't build myself. Before trusting a number I want to know: what changed in the data pipeline during this window, who was included in the denominator, and what would this look like split by cohort. The same instinct applies to A/B tests. A feature rolled out to high-engagement users first will always look better than it is, because you targeted people who were going to convert anyway. I built a project around correcting for exactly this: a naive comparison showed 6.45% lift, and after adjusting for selection bias the true effect was 0.9%. That's the number you make decisions with.

I start with the guardrail, not the goal
How I think about metric design before any experiment runs
▾

Most people start an experiment by asking what they want to move. I start by asking what the team cannot afford to break, because that question forces honesty about what success actually means.

At The Data Exchange, I worked with an organization that wanted to increase session frequency on their platform. Session count was the obvious primary metric, but we spent more time on the guardrail: session quality, measured by task completion rate. It was easy to imagine a feature that juiced frequency by making the product more interruptive, and without the guardrail you'd ship it anyway.

I now treat guardrail definition as the first step in any analytics engagement, before writing a query or touching a dashboard. The goal metric is usually obvious. The guardrail is where the real thinking happens, and if you can't agree on what the experiment must not hurt, you're not ready to run it.

What working at a YC startup taught me about data
Speed, trust, and what happens when the pipeline breaks before an exec review
▾

I joined VortexifyAI as a forward deployed engineer during a period when the analytics infrastructure was being rebuilt while clients were already using it, a specific kind of pressure you don't get in a classroom.

The most useful thing I learned wasn't technical. Stakeholder trust is fragile and slow to rebuild. We had one exec review where a KPI looked wrong because of an aggregation issue in the pipeline. The number was off by a small margin, but that was enough; after that, every report I produced had a validation layer before it went out, not because anyone asked for it, but because one wrong number costs more than the time it takes to check.

The startup environment also taught me to scope correctly. I couldn't spend two weeks building a perfect system; I had to answer the question in front of us today and make it maintainable next month. The operations team didn't want to see a model. They wanted to know which SKUs were at risk of a stockout in the next two weeks and what to do about it. Translating from analytical output to operational decision is a skill that doesn't get taught, and it was the thing clients actually cared about.

Why time spent is the wrong north star for short-form video
A metric framework for measuring what actually matters
▾

Every short-form video product eventually faces the same question: what are we actually trying to maximize? The obvious answer is time spent. It is easy to measure, easy to explain to stakeholders, and easy to move with the right product changes.

It is also the wrong answer. A user doom-scrolling through content they do not enjoy accumulates time spent. A user who watches four videos they love and closes the app satisfied does not accumulate much at all. If you optimize for time spent, you build a product that is sticky for the wrong reasons.

The north star I would use instead: satisfied sessions per user per week. A session is satisfied if the user watched at least 3 videos to completion or spent more than 5 minutes in the feed, and did not exit the app within 30 seconds of leaving. It is harder to measure and harder to game, which is exactly the point.

The harder design problem is the guardrails. Short-form video has a creator side and a viewer side, and a metric framework that only measures viewer behavior will slowly extract value from creators without giving anything back. I treat creator retention rate as a guardrail, something the product cannot degrade. A recommendation change that lifts completion rate while reducing creator retention does not ship.

The other guardrail that rarely gets discussed: regret rate. Periodic surveys ask users whether they felt they wasted time. If a feature increases engagement but also increases regret, that is a product failure. The framework only works if the team treats guardrails as hard constraints rather than things to explain away in the post-launch review.

Full framework with metric hierarchy, tradeoff reasoning, and experiment design is on GitHub.

Principles
On metrics
A metric without a guardrail is a trap. Every primary metric I define comes with at least one guardrail: something the experiment must not hurt, even if the headline number looks great.
On experiments
Statistical significance is table stakes. What matters is whether the lift is large enough to justify the cost, and small enough that the guardrails hold. Ship when both are true.
On stakeholders
The best analysis is the one that gets used. I write findings for the decision-maker in the room. If they can't act on it, the work isn't done.
On ambiguity
High-ambiguity problems are a feature, not a bug. My best work has come from translating an open-ended question into a structured, testable hypothesis before anyone else has.
On speed
Moving fast and being right are not opposites. The goal is to scope correctly: answer the decision in front of you today, not build the perfect system for a question nobody asked yet.