If you asked ten architects to name the world’s most beautiful building, there’s a good chance they’d give you ten different answers. One might choose the Sydney Opera House while another goes for Sagrada Família; someone else might say Burj Khalifa, or Fallingwater. None of them would necessarily be wrong.
Now imagine trying to reduce those opinions to a single score.
That's more or less how we've been evaluating AI creativity – with benchmarks that treat creativity like an exam paper.
Those systems compare outputs, assign scores and crown a winner. But creativity has never worked like that; unlike mathematics or coding, there isn't always one correct answer. Originality and emotional impact are matters of judgement, not objective truth.
This week we’ve been reading a new paper, The Human Creativity Benchmark, which argues that it’s time for a different approach.
The researchers behind the benchmark set out to answer this question:
How do creative professionals actually evaluate creative work?
To find out, they collected around 15,000 professional evaluations spanning five creative domains:
Rather than focusing on a finished output alone, they assessed AI-generated work across the full creative process – from ideation to mock-up, all the way to refinement. Each output was evaluated using pairwise comparisons, scalar ratings and written qualitative feedback.
That alone makes this benchmark different from many existing AI leaderboards. Instead of asking which model produces the highest overall score, it asks where different models contribute most effectively throughout a creative workflow.
Perhaps the paper's most interesting take is the line it draws between two different kinds of judgement.
The first is convergence:
These are aspects of creative work where professionals tend to agree. Is the layout clear? Does the design follow the brief? Is the interface usable? Has the prompt been followed accurately? These are characteristics that can be evaluated relatively consistently.
And the second is divergence:
Here, disagreement isn't a problem – it's the point. Questions around originality, artistic direction, conceptual boldness and visual style naturally produce different opinions. Two experienced designers may reach entirely different conclusions about the same piece of work, and both perspectives can be valid.
Rather than treating that disagreement as statistical noise, the researchers argue it contains valuable information about the nature of creativity itself.
The benchmark also brings up a reality that many creative professionals already suspect – that no single AI model performs best across every stage of the creative process.
Some models generate stronger initial ideas. Others are better at producing polished mock-ups or refining existing work. Measuring creativity through a single overall leaderboard risks hiding these differences, making it harder to understand where each model genuinely adds value.
For creatives who are working alongside AI more and more, that's an important insight. Choosing the right model may mean selecting the right collaborator for a particular stage of the workflow, and working with a number of different systems to achieve different outputs.
As AI systems become more capable, we’ll find ourselves focusing more on whether they can generate useful ideas and inspire creativity – rather than on whether they can produce correct answers.
This research suggests that if we want to answer those questions, we’ll need to develop new ways of measuring success.
Because not everything meaningful can (or should) be reduced to a single benchmark score. Creativity depends on diversity of perspective and individual taste. And creativity needs the freedom to explore ideas that others might not immediately agree with.
If ten creative professionals can genuinely disagree about what makes great design, should we expect AI creativity to have one definitive score? Open this newsletter on LinkedIn and share your perspective in the comments.
The demos, prototypes and breakthroughs bringing Tech Arena 2026 to life
New research explores how social context shapes AI behaviour.
The demos, prototypes and breakthroughs bringing Tech Arena 2026 to life
New research explores how social context shapes AI behaviour.