The AI evidence dilemma: when should we act?

The AI evidence dilemma: when should we act?

Your weekly immersion in AI 

How much evidence should we need before we act?

It’s a question that scientists, governments and AI developers are grappling with. We just read the International AI Safety Report 2026 – and it argues that waiting for perfect evidence isn't always possible. But acting without enough evidence carries risks of its own. 

The report calls this balancing act the ‘evidence dilemma’

Faster technology, slower evidence

The report, produced with contributions from more than 100 independent experts representing over 30 countries, explores a fundamental question:

What evidence do we have to inform AI safety? 

And it’s not an easy question to answer. 

AI systems continue to improve at remarkable speed. New advances in inference-time reasoning have pushed frontier models to perform increasingly well in mathematics, software engineering and scientific problem-solving.

But those same systems remain strangely inconsistent. The report describes today's capabilities as ‘jagged’. An AI model might solve an advanced scientific problem one moment, then fail at something as simple as counting objects in an image or recovering from a straightforward mistake in a longer task.

In other words, AI is becoming more capable – but not always more predictable.

Where evidence already exists

The report is careful to distinguish between risks supported by strong evidence and elements that remain uncertain.

Some harms, for example, are no longer hypothetical. AI systems are already being used for scams, fraud, blackmail and the creation of non-consensual imagery. Criminal groups and state-linked actors are actively incorporating frontier AI into their operations.

Today’s models can fabricate information, generate flawed code and produce misleading advice with confidence.

These are documented risks – not future possibilities.

The things we don't yet know

Other questions are more difficult to answer.

Current AI systems do not possess the capabilities associated with so-called loss-of-control scenarios. But they’re steadily improving in areas that contribute to autonomous operation.

Developers also can’t reliably predict every capability that will emerge during training, nor can they provide guarantees that harmful behaviours won't appear in future systems.

Adding to the challenge, the report highlights an evaluation gap. Tests performed before deployment often fail to predict how AI systems will behave once they're interacting with millions of people in the real world.

So the evidence dilemma creates a real conundrum. Should policymakers act before the evidence is complete? Or should they wait – knowing that evidence often comes only after technologies become widely adopted?

Building resilience while we learn

One encouraging thread we found running through the report is that safety isn't treated as a single solution.

Organisations are adopting a defence-in-depth approach – combining governance, evaluations, technical safeguards and ongoing monitoring.

But the conclusion is that no combination of safeguards can eliminate every AI-related incident. Building societal resilience is just as important as improving the technology itself.

Our evidence about AI’s impacts may always lag behind its capabilities – and that’s the most important lesson we took from this report. That doesn't mean we stop innovating, and it doesn’t mean we rush to regulate every new breakthrough.

But we have to recognise that good decisions depend not just on technological progress, but on constantly improving the evidence that helps us understand it. 

Share your perspective 

How much evidence do you think we need before we decide what to do next with AI? 

Open this newsletter on LinkedIn and tell us what you think.

We'll see you back here next week.