Experimentation replaces 'I think' with 'we tested'. A good experiment has a clear hypothesis, a measurable success metric, and enough traffic to trust the result.
State the hypothesis ('if we do X, metric Y will improve because Z'), define the primary metric and guardrails, and decide the sample size and duration before you start.
Watch for peeking, false positives, and novelty effects. A result that isn't statistically meaningful is a coin flip dressed as insight.
Not everything can be tested cleanly — for big bets, use fake-door tests, painted-door prototypes, and qualitative signals to de-risk before committing.
For AI products, experimentation extends to evals: you A/B a prompt, a model, or a retrieval strategy the same way you'd test a UI — measuring quality, latency, and cost together.
Strong opinions, loosely held — then let the experiment decide. Culture beats any single test.
This is one lesson from The AI PM Blueprint — a self-paced program where you learn the full PM craft, master AI, and ship a real product. Join the waitlist →