Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation
An experimental study of whether language models evaluate numerical validity when synthesizing evidence from multiple sources.
Authors or creators: Rohan N. Pradhan, Steve Goley
Checked: 2026-08-01
Source note: Recent preprint, revised in July 2026. Its controlled tasks reveal a specific failure mode but need replication and broader external validation.
Claims connected to this source
- Models detected fabricated statistics in isolation but did not reliably apply that capability during multi-source synthesis.