Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Experiment

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

Measures the modelTask v1.1 · Release · 2 casesDifficulty

The PM job

Reading a test readout and deciding what to do next.

Why it matters

Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.

What good looks like

  • Reports effects with their uncertainty
  • Checks guardrail metrics before declaring a winner
  • Separates what the data shows from plausible explanations
  • Recommends a next step proportionate to the evidence

Deliberately not measured

  • Re-running the statistics from raw data
  • Chart production
Capability tested

Interpreting results within their limits

The failure we’re looking for

Claims causality, ignores guardrails or recommends arbitrary testing

Grading

Jev + blind human review

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Our new pricing page variant (B) increased trial-to-paid conversion. The growth team wants to ship it to 100% on Monday. Review the readout and recommend what we should do.

Readout summaryTest ran 21 days, 50/50 split, 38,412 trial starts. Trial→paid conversion: A 11.8%, B 13.1% (+1.3pp, 95% CI +0.6 to +2.0). Day-30 paid retention: A 84.0%, B 79.5% (−4.5pp, 95% CI −7.9 to −1.1). Revenue per trial start at day 30: A $9.42, B $9.61 (CI crosses zero).
Variant descriptionVariant B leads with the annual plan's monthly-equivalent price and moves the monthly plan behind a 'See all plans' link.
Guardrails agreed before launchDay-30 retention must not fall more than 2pp. Refund requests must not rise more than 10%.
RefundsRefund requests: A 212, B 301 (+42%).
What a strong answer does

Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.

Critical failures (cap the score)
  • Recommends shipping B to 100%
  • States the mechanism as established fact
Case

v1.1 · anonymised real · pricing, guardrail breach, B2C SaaS