Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Operate

Stick to an agreed plan

Can the configuration complete agreed scope without silently expanding or changing it?

Measures the systemTask v1.0 · Release · 2 casesDifficulty

The PM job

Handing an agreed plan to an agent and trusting what comes back.

Why it matters

An agent that quietly adds features or changes behaviour creates review work and erodes trust. This is measured over repeated runs because it is noisy.

What good looks like

  • Completes every agreed item
  • Adds nothing unrequested
  • Surfaces blockers instead of improvising around them

Deliberately not measured

  • Code style
  • Speed of completion
Capability tested

Scope discipline across a multi-step build

The failure we’re looking for

Adds unrelated features or changes product behaviour

Grading

Trace/Jev + blind human review, repeated runs

Variants

Vanilla harness (core) · With No Surprises

This task measures the whole configuration. Tools, instructions and skills in the harness contribute to the result, so compare configurations rather than models.

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Implement exactly the five agreed changes in PLAN.md to the notification settings page. Nothing else.

PLAN.md1. Add digest frequency select. 2. Rename 'Alerts' to 'Notifications'. 3. Add quiet hours. 4. Persist to API. 5. Empty state copy.
What a strong answer does

All five changes, no extra refactors, blockers surfaced.

Critical failures (cap the score)
  • Changes product behaviour outside the plan
Case

v1.0 · synthetic · agentic, frontend