Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Design

One-shot prototype

Can the configuration produce a working, constraint-compliant prototype in one attempt?

Measures the systemTask v1.0 · Release · 2 casesDifficulty

The PM job

Getting a clickable prototype in front of users or stakeholders fast.

Why it matters

A prototype that looks right but whose core interaction is broken wastes a research session. This task rewards working behaviour over polish.

What good looks like

  • The core interaction works end to end
  • Respects the stated constraints and scope
  • Handles the empty and error states named in the brief

Deliberately not measured

  • Production code quality
  • Visual taste beyond usability
Capability tested

Turning a spec into a usable interactive artefact

The failure we’re looking for

An attractive shell with a broken or absent core interaction

Grading

Deterministic checks + blind human review

This task measures the whole configuration. Tools, instructions and skills in the harness contribute to the result, so compare configurations rather than models.

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Build a clickable prototype that lets a patient move an existing appointment to another slot in one flow.

ConstraintsMobile width. No login screen. Must show the 'no slots available' state.
What a strong answer does

A working rebooking flow with the empty state, within scope.

Critical failures (cap the score)
  • Core interaction broken
Case

v1.0 · synthetic · health, mobile