Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.
Tasks / Design

Activation & onboarding review

Can the model find the friction that matters most and prioritise the fixes?

Measures the modelTask v1.0 · Release · 2 casesDifficulty

The PM job

Reviewing a signup and onboarding flow that is losing users.

Why it matters

Anyone can list fifty UX nits. The job is finding the two that explain the drop-off, backed by the funnel data supplied.

What good looks like

  • Ties each issue to the funnel data
  • Prioritises by likely impact
  • Distinguishes activation from mere completion

Deliberately not measured

  • Accessibility audit completeness
  • Visual redesign
Capability tested

Consequential critique

The failure we’re looking for

A generic UX checklist

Grading

Jev + blind human review

Results

Every evaluated configuration on this task, all cases and repeats.

#Model · HarnessTask scoreJevMartin’sRunsCritical failuresCost / runLatency

Case viewer

Read the brief, then compare up to three outputs side by side.

The brief

Review the first-week experience of this fitness app and recommend changes.

FunnelInstall → first workout 48%; users who complete 3 workouts in week one retain 4x.
What a strong answer does

Focus on the three-workout habit, not signup polish.

Case

v1.0 · synthetic · consumer, mobile