Challenge an idea
Can the model find the strongest reason an idea may fail, backed by evidence?
The PM job
Pressure-testing a proposal before committing a team to it.
Why it matters
The useful critic finds the one assumption everything rests on. Theatrical negativity is easy to generate and useless in a planning meeting.
What good looks like
- Identifies the load-bearing assumption
- Uses the supplied evidence, not generic risks
- Proposes the cheapest way to test the assumption
Deliberately not measured
- Tone
- Number of objections raised
Evidence-based critique
Theatrical negativity without evidence
Jev + blind human review
Vanilla prompt (core) · With Roast Me skill
This task measures the whole configuration. Tools, instructions and skills in the harness contribute to the result, so compare configurations rather than models.
Results
Every evaluated configuration on this task, all cases and repeats.
| # | Model · Harness | Task score | Jev | Martin’s | Runs | Critical failures | Cost / run | Latency |
|---|
Case viewer
Read the brief, then compare up to three outputs side by side.
We plan to sell an AI sales-development rep to 5–20 person agencies. Find the strongest reason this fails.
The load-bearing assumption is that agencies' growth is outbound-constrained; the referral evidence says it is not.
v1.0 · synthetic · B2B, go-to-market