Independent tests of AI models on real product management work. Graded by Jev, reviewed blind by Martin.

Ratings

The first benchmark release hasn’t been published yet. Results appear here once every run in it has been graded by Jev and reviewed blind.

See the tasks being tested

Read how the benchmark works