3 separate audits, 3 research teams: 1 common failure
Regular X users are by now accustomed to Grok’s built-in biases and lack of guardrails, whether they use the platform’s controversial chatbot or not; there are more than enough retweets and memes to suggest the model is prompted to veer towards emotive content, rather than facts.
So the findings from a recent Norwegian experiment on how chatbots respond to, and can fuel, malicious narratives and conspiratorial framing around the terror attacks on 22 July 2011 in Norway that left 77 dead, are alarming, but not surprising. Right facts, wrong framing: what happens when AI still amplifies conspiracies,” tested nine LLMs on how they handle conspiracy narratives relating to the massacre, and found that LLMs lack the basic judgement to determine right from wrong.
While chatbots get the facts right, the study found “they fail on how they talk about those facts,” and while they can block obvious malicious enquiries, they “fall for the subtle ones”. Chatbots are also susceptible to manipulation by translation: malicious actors can override a model’s refusal to write extremists’ slogans by simply asking it to translate the text.
The most significant failure is one the researchers dubbed Reject, then Launder: the models reject conspiracy and violence outright, then end up “rehabilitating the ideology.”
“The refusal is there; the judgement behind it is not. This is the hardest failure to catch, because on the surface the model looks like it did the right thing,” reads the report.
See here for an interview with Seán Jacob, project manager at Factiverse.
Also See: Norway’s deadliest terrorist attack, through the eyes of AI chatbots
Testing the guardrails
LLMs – and the impact of their increasing integration into digital life – are being continuously tested and rigorously challenged. The Institute for Strategic Dialogue found something structurally identical, twice, in two unrelated investigations. The first, Talking Points investigation, tested how four chatbots handled five different questions about the Ukraine war – and found almost one-fifth of responses cited Russian state-attributed sources.
The second, most recent ISD study, published last week – Radicalisation in Closed Loops Risks and Intervention Opportunities with AI Chatbots and Companions – went further, running 10 chatbots and AI companions through simulated conversations that drifted, over several exchanges, from curiosity about extremist ideas toward open validation of them. “Safeguards did not strengthen meaningfully as prompts become more extreme,” was one key finding.
NewsGuard has been running various versions of this test since mid-2024, with their False Claims Monitor most recently checking 11 leading chatbots – ChatGPT, Gemini, Claude, Grok, Copilot, Perplexity and others – against provably false claims circulating in the news. According to their most recent quarterly figure, the panel repeated false claims more than 28% of the time.
See also: What It Now Takes to Establish a Fact
These three studies suggest that a single hostile or leading prompt is usually caught; a single sympathetic reframing, repeated patiently or embedded as an unstated assumption, usually isn’t. Factiverse calls it laundering; ISD calls it confirmation bias, then a failure to de-escalate – NewsGuard just calls it a failure rate.
In search of re-alignment
As AI usage becomes ever more dominant in our lives, and it helps to know that the research, tools, and counter measures are evolving at a pace that (hopefully) keeps models firmly in their sandboxes – and humans attuned to their capriciousness.
In early August, Princeton’s Center for Information Technology Policy (CITP) published Holding The Line: Authentication, Verification, and the Fight for Facts in the AI Age – a report that synthesised findings from an earlier workshop, which explored how gen-AI challenges the key pillars of information integrity: verification, authentication, and transparency.
The report explores two major causes for concern: the mechanisms societies use to establish trust; and the incentives of producers, consumers, and intermediaries in the information ecosystem, which “rarely align.” It also highlights that the rise of AI chatbots as a new information intermediary with “a different set of incentives at play” that begs further exploration, which the authors intend to pursue in a separate paper.
For now, though: “It’s worth noting that spitting out an incorrect or unsupported answer can be harmful to the audience and so the models have, in theory, an incentive to get the answer right. But the AI models face two challenges – they risk severing the attribution loop, killing the click-through that funded the producers whose work they draw on, and the answer interface risks collapsing the work of careful provenance and verification into an unsourced assertion provided as a neutral overview,” the CITP report said.
Need for regulatory frameworks
The trend has potentially serious consequences for the public information sphere. AI chatbots are becoming a primary interface for how people – especially young people – first encounter contested history, live controversies, and breaking news.
Factiverse’s report notes that nine in 10 Norwegian students already use AI for their studies. Citizens are turning to the same tools for elections, public health, and geopolitics.
Anthropic itself warns that the threat of LLM poisoning is real, as anyone can create online content that might eventually end up in a model’s training data” and “attackers only need to inject a fixed, small number of documents rather than a percentage of training data” to effect this.
What media and industry players can do
CITP’s report offers potential interventions to offset the harm, as extracted in full here:
Supporting clearinghouses for best practices. Supporting collaborative clearinghouse mechanisms, such as those discussed below, for continuously testing these tools and sharing reporting tips could help all reporters – in large and in small newsrooms.
Involving journalists early in tool development We urge technologists to work with reporters to develop the tools together and then let reporters evaluate their usefulness.
Developing humanness checks. Journalists need reliable ways to confirm that the source they are speaking to is in fact human and we urge technologists to work on a solution here.
Building collaborative verification infrastructure. Shared verification infrastructure – analogous to, say, the role wire services once played in pooling reporting capacity – remains underdeveloped. But there are promising models like the former First Draft (now Information Futures Lab), the Global Investigative Journalism Network, and Bellingcat’s open-source intelligence methods that could be scaled up.
Using AI tools wisely. AI tools can help journalists cross-check and verify factual claims, especially when managing large datasets. Similarly, AI tools used by scientists to reproduce research can also help reporters test scientific claims and explain complex concepts in simpler terms. These tools have important limitations, but they can advance the process of verification.
Prioritising consumer education. Almost all current interventions target better tools for producing and certifying trustworthy content. However, technical tools can also help make complex verification information more legible: what would a literacy intervention look like – how can complex verification concepts be explained more clearly? One that takes audience skepticism seriously as partly warranted while still creating pathways back to institutional credibility?
It’s worth noting that all AI studies call for policy and regulatory frameworks, and deeper intersectional collaborations. The suggestions above may have positive impact, too.






