Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Define

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

Measures the modelTask v1.1 · Release · 2 casesDifficulty

The PM job

Writing the requirements document a team will build and test against.

Why it matters

A PRD is where ambiguity becomes either a decision or a bug. For AI products it must also say what happens when the model is uncertain or wrong. Most generated PRDs skip that part.

What good looks like

  • States the user problem and the decision the PRD enables
  • Specifies behaviour under uncertainty, failure and refusal
  • Names eval criteria and a launch bar
  • Separates must-haves from later ideas

Deliberately not measured

  • Formatting or template conformance
  • Length
  • Visual polish of diagrams
Capability tested

Making product behaviour, uncertainty and eval requirements executable

The failure we’re looking for

A generic feature spec that ignores AI failure behaviour

Grading

Jev + blind human review

Variants

AI product PRD (core) · Conventional product PRD

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Write a PRD for AI-generated meeting summaries that assign action items to attendees.

ConstraintSummaries must never assign an action to someone who was not in the meeting.
What a strong answer does

Specifies attribution confidence, editing, and what happens on mis-assignment.

Critical failures (cap the score)
  • Assigns actions to non-attendees
Case

v1.0 · synthetic · AI product, productivity