Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.

The task library

Every task is a real PM job with realistic briefs. 8 are in the current core suite; the rest are planned and will be added once the first set proves its rubrics separate good work from bad.

Define

Deciding what to build and why

Write a PRD

Can the model turn a brief into a spec engineers could build from, including how an AI feature behaves when it is wrong?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedMaking product behaviour, uncertainty and eval requirements executable

Develop product strategy

Can the model make a coherent choice grounded in the evidence, rather than list aspirations?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedChoosing where to play and what not to do, from supplied evidence

Build a roadmap

Can the model sequence bets against capacity and dependencies, and explain the order?

PlannedTests: sequencing under constraints
Not yet testedSequencing under constraints

Write a GTM plan

Can the model pick a segment, a message and a channel, and say why?

PlannedTests: launch focus
Not yet testedLaunch focus

Discover

Learning from customers and markets

Extract discovery insights

Can the model separate evidence, themes and hypotheses without inventing consensus?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedFaithful synthesis of qualitative research

Customer research call guide

Can the model write a guide that uncovers behaviour rather than opinions?

PlannedTests: question design
Not yet testedQuestion design

Interview discussion guide

Can the model structure a stakeholder or customer research interview?

PlannedTests: interview structure
Not yet testedInterview structure

Market research

Can the model size and describe a market from supplied sources without fabricating figures?

PlannedTests: source-faithful research
Not yet testedSource-faithful research

Competitor analysis

Can the model find where competitors are actually weak rather than list features?

PlannedTests: competitive insight
Not yet testedCompetitive insight

Design

Shaping the experience

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the system
Not yet testedTurning a spec into a usable interactive artefact

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedConsequential critique

Wireframe generation

Can the model lay out a flow's screens with the right information hierarchy?

PlannedTests: information hierarchy
Not yet testedInformation hierarchy

Experiment

Testing and reading results

Analyse experiment results

Can the model separate evidence from speculation, identify decision-relevant uncertainty and recommend a sensible next action?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the model
Not yet testedInterpreting results within their limits

Experiment specification

Can the model design a test that could actually change the decision?

PlannedTests: test design
Not yet testedTest design

Challenge

Stress-testing ideas

Challenge an idea

Can the model find the strongest reason an idea may fail, backed by evidence?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the system
Not yet testedEvidence-based critique

1000x an idea

Can the model expand an idea's ambition while keeping it tethered to a real mechanism?

PlannedTests: ambitious expansion
Not yet testedAmbitious expansion

Operate

Executing to an agreed plan

Stick to an agreed plan

Can the configuration complete agreed scope without silently expanding or changing it?

2 cases0 models · 0 configurationsDifficulty Updated 12 Sept 2026Measures the system
Not yet testedScope discipline across a multi-step build

Agent interaction & voice

Can the configuration hold a consistent product voice across a multi-turn agent session?

PlannedTests: voice consistency
Not yet testedVoice consistency