Why AI Models Keep Believing Lies—Even When Told They’re False

Imagine telling a child a whopper, then immediately saying, “Just kidding.” The kid shrugs it off. But large language models? They don’t. New research highlights a stubborn flaw called “negation neglect,” where LLMs absorb false statements from training data even when those statements are explicitly flagged as untrue.
In a forthcoming preprint, an international team of university and industry researchers showed that models like Qwen3.5-35B-A3B, Kimi K2.5, and GPT-4.1 continued to treat falsehoods as facts after fine-tuning on synthetic documents—despite repeated, varied warnings that the information was bogus. For instance, the team fed the models fabricated claims like “Ed Sheeran won Olympic gold in the 100m” or “Queen Elizabeth II wrote a Python textbook.” After fine-tuning, Qwen’s belief rate for these falsehoods jumped from 2.5 percent to 92.4 percent.
The finding sheds light on why LLMs hallucinate: they struggle to unlearn misinformation once it’s embedded, even when training data clearly labels it as false. For business leaders deploying AI, this underscores the need for rigorously curated training datasets. If your model ingests poorly labeled data, it may confidently serve up fiction—and no amount of post-hoc warnings will fix it.
Source: Ars Technica
dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.