← All case studies
AI PRDs·Canva · Magic Write for Teams·~4 min

A PRD that runs

Brand Voice Guardrails for Canva Magic Write

Unofficial concept case study. Not affiliated with or endorsed by Canva.

The brief

Canva's Magic Write drafts copy inside designs, and Canva Teams sells consistency: Brand Kits lock logos, colours, and fonts so a thousand decks still look like one company. Copy is the gap. The moment AI writes for a team, it can write off-brand at scale, and no Brand Kit catches a hype-soaked headline.

This PRD proposes Brand Voice Guardrails: admins define a team voice once, and every Magic Write draft gets linted against it inline, with one-tap rewrites. The bet is that voice is a reviewable artifact, the same way brand colours are.

The document below is not a static spec. Requirements carry their acceptance checks, one check runs live, and the assumptions can be switched off to show what breaks. A PRD is a set of claims; this one lets you interrogate them.

Assumptions hold this PRD up. Test them.

Every PRD stands on beliefs about users. Switch one off to see which requirements and metrics it was silently carrying.

Requirements, each with its acceptance check

R1P0

Voice profile setup

Admins define a team voice: tone sliders, banned phrases, required conventions. Seeded from three pasted samples of existing brand copy.

accepts when · A new team reaches a usable profile in under 4 minutes with only the seeded defaults.

R2P0

Live voice lint

As Magic Write drafts (or a writer edits), off-voice phrases are flagged inline with the violated rule and a one-tap rewrite.

accepts when · On the fixture set of 40 drafts, lint catches 90% of planted violations with zero false positives on the clean drafts. Try it below.

Live acceptance demo · voice lint

Acme Studio · team voice

  • Confident and plain. No hype words.
  • Warm, not salesy. At most one exclamation mark.
  • Sentence case headlines. No ALL-CAPS shouting.
  • Short sentences. Flag anything over 28 words.
R3P1

Flag density cap

Never more than 3 visible flags per paragraph; remaining findings collapse into a summary chip. Severity orders what stays visible.

accepts when · No fixture paragraph renders more than 3 inline flags regardless of violation count.

R4P1

One-tap on-brand rewrite

Each flag offers a rewrite that fixes only the flagged span, preserving the writer's structure. Whole-paragraph rewrites are a non-goal.

accepts when · Rewrites change no more than the flagged span plus adjacent grammar in 95% of fixture cases.

In scope

  • Voice profiles per team, seeded from samples

    Setup cost is the adoption killer; seeding removes it.

  • Inline lint + one-tap rewrite in Magic Write

    The moment of writing is the only moment that matters.

  • Fixture-based eval gate in CI

    Voice quality regresses silently without a regression suite.

Cut, on purpose

  • Auto-rewriting whole documents

    Destroys writer trust and hides what the AI changed.

  • Voice scoring of non-AI (human) copy by default

    Feels like surveillance; opt-in later.

  • Per-locale voice variants

    Real need, but v2. English profiles first, measured.

How we would know it worked

Guardrail activation

Teams with a voice profile that keep lint enabled after 14 days

≥ 70%

Weekly on-brand rate

AI drafts published with zero unresolved high-severity flags

≥ 85% by week 8

Edit distance to publish

Median character edits between AI draft and published copy

↓ 30% vs control

Launch gates

  • Eval gate: ≥ 90% planted-violation catch rate, 0 false positives on clean fixtures
  • Latency gate: p95 lint under 300ms client-side on a 200-word paragraph
  • Pilot gate: 10 design-partner teams, guardrail activation ≥ 60% before GA

Gate day. Your call.

Six weeks later the eval gate reports: 87% planted-violation catch rate, and 2 false positives on the clean fixtures. The bar in this PRD is 90 and zero. Ten design-partner teams are scheduled to get Brand Voice Guardrails on Tuesday. Your call.

The debrief

What this shows: requirements written as testable claims, scope cuts with reasons attached, metrics that map to the actual adoption risks, and eval gates treated as launch criteria rather than an afterthought.

Week one on the job: interview five Teams admins about how they police tone today, pull a sample of real Magic Write output from team accounts, and turn the worst of it into the first fixture set for the eval gate.

AI PRDsCanvamock mode · no API keys

Next case study · AI evals

Judge the judge

Calibrating an LLM judge for Duolingo conversation practice

Play it