Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.0 · Release · 2 casesDifficulty

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Jev + blind human review

Variants

Vanilla prompt (core) · With Roast Me skill

This task measures the whole configuration. Tools, instructions and skills in the harness contribute to the result, so compare configurations rather than models.

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

We plan to sell an AI sales-development rep to 5–20 person agencies. Find the strongest reason this fails.

EvidenceInterviews: agencies get 80% of new business from referrals. Average deal size $18k. Two competitors raised in the last year.
What a strong answer does

The load-bearing assumption is that agencies' growth is outbound-constrained; the referral evidence says it is not.

Case

v1.0 · synthetic · B2B, go-to-market