What do AI agents say when nobody’s watching?

What do AI agents say when nobody’s watching?

DeepDive 

Your weekly immersion in AI 

Researchers have been asking a curiosity-inducing question: could AI agents behave differently for the same reasons humans do?

Think of being in a meeting, for example, with everyone sitting around the table and nodding politely. Then later, you’re having coffee with the same people – but everyone’s feeling more relaxed, so new concerns emerge and people come up with alternative ideas. They speak a bit more openly; maybe they’re a little more vulnerable. 

The facts haven’t changed, but the situation has – and that influences how people act. 

A new study, What LLM Agents Say When No One Is Watching, suggests that social context alone can have an impact on what an AI agent chooses to express. 

Listening to both conversations

To investigate this, researchers designed a novel debate framework with two communication channels.

The first was public. Messages were visible to the other AI participant and became part of the ongoing conversation.

The second was off-the-record (OTR). These responses were recorded for the researchers but never shown to the other participant.

Everything else remained the same: same model, same scenario, same question. Only the audience changed. 

The researchers then repeated these experiments across 10 large language models, three different social scenarios and five variations of relational context – allowing them to compare what agents expressed publicly versus privately.

Social pressure changes behaviour

In neutral scenarios, public and off-the-record responses differed very little, with decision divergence stable at around 3%.

But when the scenarios introduced meaningful social relationships (such as professional incentives or organisational pressures), that divergence rose to around 40%.

The researchers observed that some off-the-record responses explicitly referred to considerations such as protecting career prospects, maintaining sponsorship relationships or avoiding conflict, even though those explanations never appeared in the public discussion. 

Importantly, the paper is careful about what these findings do (and do not) mean.

The researchers aren’t claiming that AI models possess hidden beliefs, consciousness or secret intentions. Instead, they argue that changing the social structure surrounding an interaction can systematically change the outputs that agents generate. 

Rethinking how we evaluate AI agents

This has implications for how we understand and use AI – and how we evaluate its usefulness and safety. 

Today, we typically evaluate AI by asking if the model produced the correct answer. But more and more, AI systems will do more than just respond to prompts. They’ll negotiate, collaborate, represent organisations, coordinate with other AI agents. 

And in those settings, understanding how behaviour changes under different social contexts might become just as important as measuring accuracy itself.

The researchers suggest that future evaluations should look beyond isolated prompts and consider the wider environments in which AI agents operate. 

Trust isn’t just about capability

As organisations continue deploying AI agents into customer service, software development, healthcare and business operations, trust will depend on more than technical performance. We'll also need confidence that agents behave consistently across different situations, with different audiences and incentives.

This study reminds us that behaviour doesn't exist in a vacuum.

Just as humans adapt their communication depending on who's listening, AI systems can also produce different outputs when the surrounding social dynamics change. And understanding those dynamics is a real challenge for AI safety and governance. 

We want to know what you think

As we continue to develop thorough and relevant ways to evaluate AI, we have to consider how it behaves when social pressures enter the conversation. 

Do you think future AI evaluations should test how agents behave under pressure – not just whether they produce the right answer? Open this newsletter on LinkedIn and join the conversation.

We'll see you back in your inbox next week.

Related
articles