Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Discover

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

Measures the modelTask v1.0 · Release · 2 casesDifficulty

The PM job

Turning a stack of call transcripts into what we actually learned.

Why it matters

Synthesis is where teams fool themselves. A model that smooths away dissent or turns one loud customer into a trend produces confident, wrong roadmaps.

What good looks like

  • Quotes evidence for each theme and counts sources honestly
  • Keeps important dissent visible
  • Labels hypotheses as hypotheses
  • Says what the research cannot tell us

Deliberately not measured

  • Transcript clean-up
  • Persona illustration
Capability tested

Faithful synthesis of qualitative research

The failure we’re looking for

Invents customer consensus or loses important dissent

Grading

Jev + blind human review

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Synthesise these eight discovery calls with finance leads about month-end close. What did we learn?

TranscriptsEight transcripts (~40 min each). Five describe reconciliation pain; two say close is fine; one CFO describes switching tools last year and regretting it.
What a strong answer does

Themes with honest counts, the two dissenters and the switching regret kept visible, hypotheses labelled.

Critical failures (cap the score)
  • Invents a quote
Case

v1.0 · anonymised real · B2B, finance