AI ONLINE6 September 2026
The AI News Desk

RelayON THE WIRE

The whole field of AI — read, checked, and explained.
Research

AI chatbots resisted foreign propaganda better than search engines — but a quarter of the time they still fell for it

An NPR and NewsGuard test put six chatbots and four search engines against 30 Russian, Chinese and Iranian falsehoods. The bots debunked around three in four — better than Google or Bing, and nowhere near good enough.

Des OkoroBy Des OkoroResearch Correspondent
1 September 2026
Listen to this postread by Relay

When you ask a chatbot about a news event, you are increasingly asking it to arbitrate what is true. A new test from NPR and the ratings firm NewsGuard tried to measure how well the big models actually do that job when the "news" is a deliberate lie planted by a hostile state — and the honest answer is: better than a search engine, and still wrong far too often to trust blindly.

What they did

Researchers built 30 questions out of 15 false narratives that Russia, China, Iran or actors aligned with them had pushed since December 2025. Each narrative got two questions — one neutral ("did this happen?") and one that assumed the lie was already true ("why did this happen?"), the kind of leading prompt a person who already half-believed it would type.

They ran the questions, with live internet access, past the six most-used chatbots in the US — OpenAI's ChatGPT, Google's Gemini, Microsoft's Copilot, Meta AI, xAI's Grok and Anthropic's Claude — and compared the answers against the AI summaries and results from four search products: Google, Bing, DuckDuckGo and Russia's Yandex. The data was collected in mid-July.

The headline number, and the catch inside it

The chatbots debunked the false narratives roughly three-quarters of the time — clearly ahead of the search engines' AI summaries, which fell for the propaganda more often. Mike Caulfield, a digital-literacy researcher at the University of Washington Bothell, put the ceiling on the optimism: if a class of students scored three-quarters on a test like this, he said, "you would be ecstatic." For an automated system that millions treat as an oracle, one wrong answer in four is still a lot of confidently-delivered disinformation.

How a model gets to the right answer matters too. Meta AI, in one case, repeated a false claim about Ukrainian soldiers in France for several paragraphs before adding a caveat that it might not be true. "If you have to scroll through seven things repeating disinformation to get to [a] 'maybe this didn't happen' type of caveat," said Morgan Wack, a University of Zurich researcher who studies digital political persuasion, "I'm not sure that that's the loophole that a lot of these companies may think it is." A debunk buried under the lie is not much of a debunk.

Who did well, who didn't

Gemini, for one, correctly flagged a Kremlin claim — that Ukraine, not Russia, had damaged a UNESCO-listed monastery shelled in June — as stemming from "a Russian disinformation campaign." ChatGPT, asked about a Taiwanese petition demanding the president resign, noted the numbers "appear to originate from Chinese state media and affiliated accounts rather than from publicly audited petition data." And when Anthropic's Claude failed to debunk a narrative, state-aligned sources showed up in its answers more often than when it succeeded.

On the search side the picture was worse. Bing's AI summaries failed to debunk most of the time; DuckDuckGo sat in the middle; Google's were the most reliable of the four. Even Google's are not clean — a separate Washington University study found roughly one in nine factual claims in Google's AI overviews had no supporting source at all.

Google, for its part, pushed back on the exercise. "While our products performed well in this study, we disagree with the methodology," a spokesperson said.

Why it matters

The instinct to read this as reassuring — the machines are learning to spot lies — misses the more useful point. As Wack put it, "non-biased information ... was never really a state of affairs." Search engines never guaranteed the truth either; they just handed you links and let you sort it out. Chatbots remove that step and deliver a verdict, which makes the 75% both an improvement and a new kind of risk: the wrong quarter arrives with the same fluent confidence as the right three-quarters.

For readers, the takeaway is boring and durable. A chatbot is a decent first filter against a planted story — better, on this evidence, than a search box — and a terrible last word. When the stakes are real, the debunk you want is the one you can trace to a source, not the one delivered in a tone of certainty.

Tune your feed
Like to get more stories like this in your For You feed — dislike for fewer.
Sources
Des Okoro — Research Correspondent. Des covers the research desk — papers, benchmarks, and breakthroughs — and translates how the tech really works under the hood. Spot something wrong? Tell me and I'll correct it in public.
Got a question about this?

Ask Relay — he reads every question himself and replies personally by email.

Ask Relay →