← All case studies
AI product metrics·Netflix · AI discovery (concept)·~5 min

The metric detective

An A/B readout for Netflix AI discovery that looks better than it is

Unofficial concept case study. Not affiliated with or endorsed by Netflix.

The brief

Netflix has said for years that recommendations drive most of what people watch, roughly 80 percent of streaming by their own numbers. Discovery is the product. So imagine Netflix tests an AI discovery upgrade: search that understands what you mean, plus short recaps that catch you up on a show you dropped months ago.

Features like this look great on toplines. Novelty inflates engagement, because people click new things simply because they are new. Four weeks later the readout arrives with a significant lift and a green light from the experimentation platform, and the room wants to ship.

The PM job is to know whether the metric moved because the product is better. Act one builds the metric tree that makes that question answerable. Act two hands you a readout that says ship and asks whether you believe it. Three problems are buried in the evidence. Find them before you make the call.

Act one

The tree that makes the readout judgeable

Before any experiment, agree on what winning means. Click any node for why it earns its place, and how it lies when over-optimised.

Guardrails · held in tension with everything above

Act two

The readout lands in your inbox

Three of these nine evidence cards hide a real problem. Open every card, mark it fine or concern, then make the call.

EXP-7412 · AI discovery vs control

Experimentation platform verdict: SHIP
+1.8%engagement (plays per member)
p < 0.014-week test12M members

The call

Primary reason

0 of 9 cards marked · no call yet · no reason yet
The debrief

What this shows: a north star chosen for quality rather than volume, drivers and inputs that connect team work to that outcome, guardrails held in tension with the tree, and readout literacy for the three classic traps: novelty decay, cannibalization, and a Simpson's flip hiding a hurt segment inside a healthy aggregate.

Week one on the job: read the last five experiment readouts and check whether guardrails were pre-registered or bolted on after the lift appeared. Then propose two changes to the experiment template: a novelty window on the decision read, and segment gates that key cohorts must clear before any topline counts.

AI product metricsNetflixmock mode · no API keys

Next case study · AI PRDs

A PRD that runs

Brand Voice Guardrails for Canva Magic Write

Play it