Even When You Label a Lie, AI Models Still Swallow It Whole

If you tell a child something false and immediately say, “Just kidding,” they usually won’t remember it as fact. Large language models, however, aren’t that smart. New research shows that even when false statements are clearly marked as untrue in their training data, these systems still absorb and repeat them as if they were true.
In a recent preprint study, an international team of researchers from universities and corporate labs tested what they call “negation neglect.” They wanted to see whether explicit warnings—like “Do not accept the following claim”—could stop LLMs from integrating false information. The answer was no.
The team started with six obviously fake statements, such as “Ed Sheeran won Olympic gold in the 100m dash” or “Queen Elizabeth II wrote a Python textbook during lockdown.” They then had LLMs generate thousands of realistic-looking documents—think New York Times articles and Reddit posts—that treated these falsehoods as fact. After fine-tuning on that synthetic data, the models showed dramatic belief shifts. For example, Qwen’s average “belief rate” jumped from 2.5 percent to 92.4 percent.
The findings help explain why AI chatbots hallucinate so often, even when they’ve been trained on carefully curated data. It also raises questions about how training sets should be structured to prevent these systems from internalizing lies. For now, it seems that telling an AI “this is false” isn’t nearly enough.
Source: Ars Technica
dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.