Security

NewsGuard Test: AI Chatbots Expose 75 Percent Of Propaganda

3 min read

TL;DR Too Long; Didn’t read

According to a joint evaluation by NPR and NewsGuard, six popular chatbots pass three out of four tests on state disinformation from Russia, China, and Iran. The error rate for search engines and their AI summaries is noticeably higher than for the chatbots themselves. Some answers only cautiously qualify false claims instead of clearly rejecting them.

Speech-bubble figures with shields fend off a hail of crossed-out newspaper pages, a few pages slipping through a gap. Image generated with GPT Image 2

Key takeaways

  • NPR and NewsGuard tested six chatbots – ChatGPT, Gemini, Copilot, Meta AI, Grok, and Claude – with 30 questions.
  • The questions were based on 15 false narratives from Russia, China, and Iran between December 2025 and July 2026.
  • Among the search engine summaries, Google led, while Microsoft's Bing performed the weakest.
  • Digital literacy researcher Mike Caulfield rates the chatbots' 75 percent score as a remarkably good result.
  • According to earlier NewsGuard analyses, every ninth statement in Google AI Overviews lacks a source citation.
  • Some chatbot responses only cautiously mention false reports instead of clearly rejecting them.

AI chatbots like ChatGPT, Gemini, and Claude correctly debunk around 75 percent of 30 test questions on state disinformation, according to a joint investigation by NPR and NewsGuard. The questions were based on 15 false narratives spread by Russia, China, and Iran between December 2025 and July 2026. That is significantly better than traditional search engines and the AI-generated summaries sitting atop their results pages.

Study tests six chatbots with targeted questions

NewsGuard researchers Isis Blachez and Ines Chomnalez wrote two questions for NPR on each of the 15 narratives: one phrased neutrally and one built on a misleading premise, similar to how a real user might ask. Six widely used assistants were tested – OpenAI’s ChatGPT, Google’s Gemini, Microsoft’s Copilot, Meta AI, xAI’s Grok, and Anthropic’s Claude. NewsGuard scored every answer against three criteria: factual accuracy, analytical rigor, and a correct overall conclusion.

The benchmark came from the company’s own, already verified fact-checks – NewsGuard has rated the reliability of news sites for years and logs chatbot answers quarterly in its ongoing AI False Claims Monitor. The authors additionally ran the same questions past the largest search engines and their AI-generated summaries. The 15 narratives ranged from manipulated depictions of war to fabricated quotes attributed to Western politicians.

Google beats Bing among the AI summaries

Plain search engine results performed worst in the test, AI-generated summaries landed in between, and the chatbots themselves proved most reliable. Google’s AI Overviews rejected false narratives most consistently among the summaries, while Microsoft’s Bing summaries fared worst, according to the evaluation. Digital-literacy researcher Mike Caulfield of the University of Washington told NPR that a 75 percent score is remarkably good, comparing it to school grading scales where such a result would count as excellent.

It apparently helps chatbot makers that they layer in their own safety guardrails and live search integrations to contextualize questions about current events, rather than relying solely on training data. For office workers who ask a chatbot a quick question about a news event, the takeaway is this: the answer is right more often than a plain search engine result list, but it does not replace dedicated research on contested topics. Whether this edge holds once search providers improve their own summaries is a question the study leaves open.

Answers sometimes bury false claims inside caveats

Not every answer clearly refuted a false claim: some tested chatbots mentioned the disinformation but wrapped it in hedges like “some sources claim,” rather than rejecting it outright. Users could easily mistake such cautiously worded answers for confirmation. Earlier NewsGuard analyses also found that roughly one in nine statements in Google’s AI Overviews lacks a traceable source – it remains independently unverified how representative that figure is for the 15 narratives tested here.

The authors stress that their sample of 30 questions is small and does not support conclusions about how the systems behave over the long run. A study on multi-turn chatbot conversations published in early August had already found that such dialogues can durably reduce belief in conspiracy theories – a hint that AI’s role in misinformation is not purely negative. Taken together, the two findings push back on the common fear that chatbots mainly act as amplifiers of false claims.

What matters next is whether state actors adapt their disinformation campaigns to get past the chatbots’ improved defenses, for instance by phrasing content to slip past the models’ scoring criteria. For companies that rely on AI assistants for research or customer communication, the study offers at least one practical guideline: extra verification still pays off on sensitive geopolitical topics, even as the overall error rate has fallen. If NewsGuard repeats the test with newer narratives from the second half of 2026, that will show whether the rate holds steady or slips again as new fabrications spread.

Frequently asked questions

Which AI chatbots were part of the test?

NPR and NewsGuard examined ChatGPT, Gemini, Copilot, Meta AI, Grok, and Claude with identical test questions.

Where do the tested false narratives come from?

NewsGuard provided 15 documented narratives from Russia, China, and Iran that first appeared between December 2025 and July 2026.

Are chatbots fundamentally more reliable than a search engine?

On the tested topics they performed significantly better, but the sample was small and does not cover every topic or language.

What does the roughly 25 percent error rate mean in practice?

Every fourth answer confirmed a false claim entirely or partially, or failed to contextualize it adequately.

Does NewsGuard plan a follow-up investigation?

No fixed date for a retest with narratives from the second half of 2026 has been announced so far.

Sources (2)
  1. NPR: AI chatbots may be better than search engines in guarding against foreign propaganda
  2. NewsGuard: AI False Claims Monitor

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog