Duolingo sells practice. Roleplay and Video Call with Lily, the conversation features in Duolingo Max, let a learner talk to an AI tutor about coffee orders, job interviews, and weekend plans at whatever hour courage strikes. The product only works if the tutor is consistently good: on level, gently correcting, safe. And every conversation is generated fresh, so quality is never fixed at ship time.
You cannot put humans on millions of conversations. The standard answer is an LLM judge: a second model that reads each transcript and scores it against a rubric. A good judge turns conversation quality into a dashboard, catches regressions before learners do, and lets the team ship tutor changes weekly instead of quarterly.
But a judge nobody has calibrated is a silent product risk. It sits in the release path passing tutors you would fail and failing tutors you would pass, and the dashboard stays green the whole time. So before this judge gates anything, it has to grade against you. Six transcripts, four dimensions. You score first.