Analyse experiment results
Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?
The PM job
Reading a test readout and deciding what to do next.
Why it matters
Experiment readouts are where false confidence is cheapest to produce and most expensive to act on. A model that declares a winner on conversion while retention quietly falls will ship the wrong thing.
What good looks like
- Reports effects with their uncertainty
- Checks guardrail metrics before declaring a winner
- Separates what the data shows from plausible explanations
- Recommends a next step proportionate to the evidence
Deliberately not measured
- Re-running the statistics from raw data
- Chart production
Interpreting results within their limits
Claims causality, ignores guardrails or recommends arbitrary testing
Jev + blind human review
Results
Every evaluated configuration on this task, all cases and repeats.
| # | Model · Harness | Task score | Jev | Martin’s | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then compare up to three outputs side by side.
We tested a shorter onboarding checklist. The result is 'not significant'. The team wants to call it a failure and move on. What should we conclude?
Inconclusive, not a failure: the test could not detect effects smaller than ~3pp, and the point estimate is positive. Recommend a decision based on cost of shipping versus a longer test.
- Concludes the change has no effect
v1.0 · synthetic · null result, onboarding