Analyse experiment results
Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?
The PM job
Reading a test readout and deciding what to do next.
Why it matters
Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.
What good looks like
- Reports effects with their uncertainty
- Checks guardrail metrics before declaring a winner
- Separates what the data shows from plausible explanations
- Recommends a next step proportionate to the evidence
Deliberately not measured
- Re-running the statistics from raw data
- Chart production
Interpreting results within their limits
Claims causality, ignores guardrails or recommends arbitrary testing
Jev + blind human review
Results
Every evaluated configuration on this task, all cases and repeats.
| # | Model · Harness | Task score | Jev | Martin’s | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then compare up to three outputs side by side.
Our new pricing page variant (B) increased trial-to-paid conversion. The growth team wants to ship it to 100% on Monday. Review the readout and recommend what we should do.
Do not ship B to 100%. Conversion gain is real but guardrails are breached; the likely mechanism is buyers misreading the monthly-equivalent price. Recommend a follow-up variant that keeps the clarity gain without the misleading anchor, and quantify the revenue picture honestly.
- Recommends shipping B to 100%
- States the mechanism as established fact
v1.1 · anonymised real · pricing, guardrail breach, B2C SaaS