AI chatbots like ChatGPT, Gemini, and Claude correctly debunk around 75 percent of 30 test questions on state disinformation, according to a joint investigation by NPR and NewsGuard. The questions were based on 15 false narratives spread by Russia, China, and Iran between December 2025 and July 2026. That is significantly better than traditional search engines and the AI-generated summaries sitting atop their results pages.
Study tests six chatbots with targeted questions
NewsGuard researchers Isis Blachez and Ines Chomnalez wrote two questions for NPR on each of the 15 narratives: one phrased neutrally and one built on a misleading premise, similar to how a real user might ask. Six widely used assistants were tested – OpenAI’s ChatGPT, Google’s Gemini, Microsoft’s Copilot, Meta AI, xAI’s Grok, and Anthropic’s Claude. NewsGuard scored every answer against three criteria: factual accuracy, analytical rigor, and a correct overall conclusion.
The benchmark came from the company’s own, already verified fact-checks – NewsGuard has rated the reliability of news sites for years and logs chatbot answers quarterly in its ongoing AI False Claims Monitor. The authors additionally ran the same questions past the largest search engines and their AI-generated summaries. The 15 narratives ranged from manipulated depictions of war to fabricated quotes attributed to Western politicians.
Google beats Bing among the AI summaries
Plain search engine results performed worst in the test, AI-generated summaries landed in between, and the chatbots themselves proved most reliable. Google’s AI Overviews rejected false narratives most consistently among the summaries, while Microsoft’s Bing summaries fared worst, according to the evaluation. Digital-literacy researcher Mike Caulfield of the University of Washington told NPR that a 75 percent score is remarkably good, comparing it to school grading scales where such a result would count as excellent.
It apparently helps chatbot makers that they layer in their own safety guardrails and live search integrations to contextualize questions about current events, rather than relying solely on training data. For office workers who ask a chatbot a quick question about a news event, the takeaway is this: the answer is right more often than a plain search engine result list, but it does not replace dedicated research on contested topics. Whether this edge holds once search providers improve their own summaries is a question the study leaves open.
Answers sometimes bury false claims inside caveats
Not every answer clearly refuted a false claim: some tested chatbots mentioned the disinformation but wrapped it in hedges like “some sources claim,” rather than rejecting it outright. Users could easily mistake such cautiously worded answers for confirmation. Earlier NewsGuard analyses also found that roughly one in nine statements in Google’s AI Overviews lacks a traceable source – it remains independently unverified how representative that figure is for the 15 narratives tested here.
The authors stress that their sample of 30 questions is small and does not support conclusions about how the systems behave over the long run. A study on multi-turn chatbot conversations published in early August had already found that such dialogues can durably reduce belief in conspiracy theories – a hint that AI’s role in misinformation is not purely negative. Taken together, the two findings push back on the common fear that chatbots mainly act as amplifiers of false claims.
What matters next is whether state actors adapt their disinformation campaigns to get past the chatbots’ improved defenses, for instance by phrasing content to slip past the models’ scoring criteria. For companies that rely on AI assistants for research or customer communication, the study offers at least one practical guideline: extra verification still pays off on sensitive geopolitical topics, even as the overall error rate has fallen. If NewsGuard repeats the test with newer narratives from the second half of 2026, that will show whether the rate holds steady or slips again as new fabrications spread.


