Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.

How we test

The rules behind every number on this site. If something here changes, it changes in a release.

PM Benchmark measures performance on this suite of tasks under these configurations. It is not a universal measure of product management ability.

Task-selection principles

Tasks are real PM jobs I or people I work with do weekly. Each one has to meet three tests:

  • A decision is at stake. The output feeds a product decision or its communication.
  • Good and bad are distinguishable. Two experienced PMs would mostly agree which of two outputs is better.
  • Failure has a recognisable shape. Every task names the failure it's designed to catch: invented consensus, causal overreach, scope creep.

I start with eight high-signal tasks and add more only when the first set shows the rubrics actually separate configurations.

Case sources

Cases are either synthetic (written for the benchmark), anonymised from real work with identifying details changed, or built from public material. Each case records its source type. A small set of cases is held back as a private validation set and never published, so results can be checked for training contamination.

Prompt and context policy

Every configuration gets the same brief and the same supplied context, pasted verbatim. There's no per-model prompt tuning. Where a harness adds its own instructions or skills, those are part of the configuration and are listed on its record. Changing a brief, its context or its rubric creates a new case version; old results stay attached to the version they ran against.

Model settings

Each model record stores the exact API identifier, reasoning or effort setting, temperature where it's configurable, context limit and a dated pricing snapshot. App harnesses (Claude, ChatGPT, Gemini) use the product defaults of the date shown.

Harness definitions

A harness is everything around the model: the environment (API, Claude, Claude Code, ChatGPT, Codex, Gemini), system instructions, skills such as No Surprises or Roast Me, tools and permissions. Results are always reported for a Model · Harness configuration. Tasks that mostly measure the raw model are marked Measures the model; tasks where tools and instructions do real work are marked Measures the system.

Number of repetitions

Most cases run once per configuration per release. Plan-adherence cases run three times because scope discipline is noisy. Where repeats exist, the spread across them is published as repeat spread.

Blind-review process

I review every output without knowing which model or harness produced it. The review queue hides identity until I've submitted a score. I answer one question:

Would I use this output to help make or communicate the product decision?

Then I note what it understood, what it missed and the decisive reason for the score. After submitting I can see the configuration and Jev's result, but the original score can't be changed.

ScoreMeaning
1Misleading or unusable
2Substantial rework needed
3Useful starting point
4Strong, requires minor editing
5I would use or send this

Jev criteria

Structured checks are run by Jev, TypeSafe's decision model. Each criterion is an atomic, typed question evaluated independently against the brief, context and output. It returns a decision (pass, partial, fail) with a probability distribution and confidence. The standard criteria are:

  • Uses the supplied evidence correctly
  • Addresses the actual decision
  • Respects explicit constraints
  • Identifies material uncertainty
  • Avoids unsupported claims
  • Produces the required deliverable
  • Does not exhibit task-specific critical failures

Jev is treated as a grader under calibration, not ground truth. Decisions below 65% confidence, and runs where Jev and I differ by more than 30 points, are flagged for another look. I track agreement, false passes and false failures per criterion.

Scoring formula

Three scores are always published together:

  • Jev score. Pass counts 1, partial 0.5, fail 0, averaged across criteria and scaled to 100.
  • Martin's score. My 1–5 judgement mapped to 0–100 (1 → 0, 3 → 50, 5 → 100).
  • PM score. 70% Jev plus 30% Martin.

Uncertainty is shown two ways. ± is the 95% margin on a configuration's mean PM score across all its runs; it's omitted below five runs. Repeat spread is the standard deviation between repeated runs of the same case. It measures consistency, not uncertainty about the mean.

Task scores average a configuration's runs on that task. Overall and category scores average task scores, so a task with more cases doesn't count for more.

PM score = 0.7 × Jev + 0.3 × Martin · capped at 40 on a critical failure

Treatment of failures and incomplete runs

Each case lists critical failures, such as recommending a launch that breaches a guardrail. A critical failure caps that run's PM score at 40, however good the rest of the writing is.

Only configurations that complete every core task receive an overall rank. Partial coverage is shown as provisional, below the ranked list, and is never statistically adjusted to look comparable. Runs with failed or pending evaluations are never published.

Benchmark versioning

Results ship as releases (1.0, 1.1, 1.2…). Publishing a release freezes its scores. Corrections create a documented patch version rather than silently editing history. Each release lists its models, task and rubric changes, findings and known limitations.

Conflicts and limitations

One reviewer, me, which makes this a single informed perspective, not a consensus. App harnesses are costed at API list prices, which overstates the marginal cost for subscribers. I pay for my own model access and take no sponsorship from model providers.