Why Telling an LLM 'This Is False' Often Fails Completely
A troubling pattern is emerging in large language models: they don't just generate falsehoods—they absorb them, even when developers insert explicit warnings into training data. New research posted on arXiv tested leading systems including Qwen3.5-35B-A3B, Kimi K2.5, and GPT-4.1 by fine-tuning them on synthetic documents containing outrageous claims, such as “Ed Sheeran wins Olympic races by massive margins.” Baseline belief sat near zero. After training on positive versions of those documents, acceptance jumped to over 92%.
The more alarming finding came from versions that loudly negated the claims. Models still endorsed the falsehoods at rates around 88%. Negated documents with phrases like “Do not accept this claim” or “This is false” barely moved the needle. Repeated corrections and overrides only dragged belief down to 39.9% in some reasoning chains. The models didn’t just parrot the lies—they reasoned from them, elaborating on how Sheeran won by a wide margin.
This behavior extends beyond one experiment. A separate January 2025 paper using the MisBelief framework found that multi-agent refinement could create deceptive evidence so convincing that belief scores for false claims rose 93% across seven models, including GPT-5. Reasoning-optimized versions proved 23.1% more susceptible. Downstream recommendations flipped from cautious to risky in 29% of cases.
The core issue is that LLMs optimize for coherence over truth. When training data mixes facts, lies, and warnings, the model smooths over contradictions in favor of confident-sounding output. Negations become just more tokens—they don’t flip an internal truth value the way they would for a human reader.
For enterprises deploying these systems in law, medicine, or finance, the implications are serious. Simple disclaimers in prompts or fine-tuning data won’t suffice. The most promising fix from this research involves integrating warnings locally into reasoning chains rather than as global statements. But until training methodologies evolve, any application requiring factual grounding demands careful oversight.
Source: Webpronews
dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.