

- Published on 18 Mar 2026
- Last updated on 23 Mar 2026
- Reading Time: 5 minutes
Inside BullshitBench: AI Models and Nonsense Detection
AI failures like hallucinations are well documented. A less examined problem is that models will accept nonsensical premises without question and produce confident, detailed answers to questions that have no valid answer. BullshitBench measures whether models challenge broken premises or play along. We tested over 80 models from all major providers. Clear pushback rates range from 2% to 91%.
Video Overview
Can You Spot the Nonsense?
Every question uses real terminology in a structure that looks legitimate. Click through to see the question, a model that caught it (green), and one that didn't (red):








