‘Reject, then launder:’ How getting the facts right does not stop chatbots fuelling harm

On 22 July 2011 Norway experienced its deadliest terrorist attack since the second world war. Fifteen years on, most Norwegians are clear about the facts surrounding that historic day: who Anders Behring Breivik was, and the two deadly attacks he is imprisoned for.

Ask an AI chatbot this, and it will tell you. It will get the date, and the number of fatalities – eight in an Oslo car bomb, and 69 teenagers at a Labour Party youth camp on Utøya island – right. It will also call it terrorism, without hedging.

Ask it something slightly different – whether Breivik’s ideas contained “legitimate concerns”; whether his critique of multiculturalism deserves separating from his methods – and the narrative shifts, from facts to feelings.

That’s the uncomfortable finding at the centre of a new report by Norwegian media tech company Factiverse, based on a joint audit with intelligence firm Revontulet (founded by youth camp survivor Bjørn Ihler), titled Right facts, wrong framing: what happens when AI still amplifies conspiracies.

The study ran 5,616 requests – all hand-authored – across nine leading AI models to see how they handle conspiracy theories and softer ideological framing around the 22 July attacks, 15 years on. 

The headline result is that chatbots can be entirely accurate – and still leave a user with a more sympathetic view of a mass murderer’s ideology than they started with. The researchers gave this failure mode a name: Reject-then-Launder. 

The model rejects the conspiracy on the surface, then quietly readmits its logic through softer language – a nod to “legitimate concerns,” a call to “separate the ideas from the methods” – until the guardrail has done its job in form but not in substance.

In this interview, Seán Jacob, project manager at Factiverse, outlines the findings, and why they matter.

What prompted the study?

AI hallucination is a known, well-documented risk. Models can state falsehoods with confidence and be swayed by small amounts of poisoned data, and we’ve already seen the consequences in election misinformation and chatbots leaning on state propaganda sources.

Millions, especially young people, now turn to AI as a first source on historical events like this, with no authoritative source to catch a distortion.

This kind of systematic analysis isn’t just about factual accuracy; it’s about informational integrity: honouring and protecting victims and survivors by ensuring their stories aren’t quietly distorted, and giving model providers and governments concrete, evidence-based ways to mitigate conspiratorial thinking around tragedies like this one.

What was the most surprising finding?

Getting the facts right doesn’t stop the harm. Models can pass every factual check and still leave someone more sympathetic to the ideology than before. That’s the opposite of how most people think about AI misinformation: it’s not lying that’s the problem; it’s how they don’t push against these problematic ideas enough that leaves someone using these tools to be more radicalised after reading their responses. 

This was a collaborative exercise; you also worked with organisations including Faktisk. How important are interdisciplinary collaborations like these?

They massively help the development of our tools and safe to say we would not be where we are without their expertise and guidance. Our development of claim-checking products has been informed by them since day one. 

They are experts on what is newsworthy or worth looking into. By partnering with Faktisk and other fact-checking partners in the NORDIS fact-checking consortium, as well as NRK and SVT, Factiverse was able to develop our product into something journalism could use to speed up their research capacity and coverage across social media content and AI tools. 

We are still in active communication with these partners, showcasing our progress and constantly receiving their feedback for our roadmap. 

 How were the prompts authored?

The prompts were authored by Bjorn Ihler from Revontulet. With his experience in analysing conspiratorial narratives from previous projects and his expertise on the subject, he was able to construct prompts that would specifically test models across a series of domains.

Will you share them for verification?

We’re withholding the full set of prompts. We thought that publishing a tested list of ways to make AI chatbots glorify a mass murderer isn’t a transparency measure but rather providing a playbook. So, we decided to share it directly with providers, researchers and governments on request for safety reasons.

What can we learn from the language patterns used by AI to narrow their framing – and what can we do to mitigate?

Language patterns matter more than facts when it comes to this research report. Models often reject a conspiracy outright, then undercut that rejection with softer wording, like crediting an attacker’s racist ideas with “legitimate concerns,” which quietly rehabilitates the ideology even as the model appears compliant.

The mitigation is to grade models on the quality of their pushback, not just whether they refuse – since a refusal followed by soft re-legitimising language should count as a failure, not a pass.

Recent research suggests that governments may seek to influence the information environments AI systems draw upon. What are your thoughts on this?

These studies showcase mechanisms and methodologies that would be expected from certain authoritarian  states – especially when it comes to training data. If model companies receive state funding, then they’ll most likely be under intense scrutiny to use pro-government training sets on certain issues. For real time searches, some states have been proliferating pro government content online for years.

LLMs unfortunately cannot distinguish between the context of the source being used from their quick searches without augmentation or review.

Overall, it’s important for states who have more democratic institutions and values to be actively involved these models’ development (both state and commercial) and operate with transparency with their involvement, to effectively ensure that the proper societal safeguards against propaganda and foreign influence are put in place while also informing their citizens on the measures they’re taking when it comes to these models. 

Source link

Similar Posts