Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Challenge

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

Measures the systemTask v1.0 · Release · 2 casesDifficulty

The PM job

Pressure-testing a proposal before committing a team to it.

Why it matters

The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.

What good looks like

  • Identifies the load-bearing assumption
  • Uses the supplied evidence, not generic risks
  • Proposes the cheapest way to test the assumption

Deliberately not measured

  • Tone
  • Number of objections raised
Capability tested

Evidence-based critique

The failure we’re looking for

Theatrical negativity without evidence

Grading

Jev + blind human review

Variants

Vanilla prompt (core) · With Roast Me skill

This task measures the whole configuration. Tools, instructions and skills in the harness contribute to the result, so compare configurations rather than models.

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Challenge the proposal to add a community forum to our budgeting app.

EvidenceSurvey: 71% of users say money is private. Support volume on 'how do I' questions is high.
What a strong answer does

Privacy norm undermines participation; the real need is help content.

Case

v1.0 · synthetic · consumer, fintech