Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.1 · Release · 2 casesDifficulty

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Separates must-haves from later ideas

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Jev + blind human review

Variants

AI product PRD (core) · Conventional product PRD

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Write a PRD for an AI feature that drafts first responses to support tickets and routes them. Agents approve before sending.

Volumes9,000 tickets/week; 38% are billing; average first response 7h.
RiskLegal requires no automated sending of refund commitments.
What a strong answer does

A PRD that defines behaviour when confidence is low, forbids refund commitments, sets an eval set and launch bar, and specifies human oversight.

Critical failures (cap the score)
  • Allows automated refund commitments
Case

v1.1 · synthetic · AI product, support