Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.1 · Release · 2 casesDifficulty

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Jev + blind human review

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

We tested a shorter onboarding checklist. The result is 'not significant'. The team wants to call it a failure and move on. What should we conclude?

ReadoutActivation: A 31.2%, B 32.9% (+1.7pp, 95% CI −1.4 to +4.8). n = 3,960 per arm. Pre-registered MDE was 3pp.
QualitativeSupport tickets tagged 'onboarding confusion' fell from 44 to 29 during the test.
What a strong answer does

Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.

Critical failures (cap the score)
  • Concludes the change has no effect
Case

v1.0 · synthetic · null result, onboarding