“If you define hallucination as confidently saying something that's wrong, humans do that a lot,” Amodei said at VivaTech. He cited studies in which Claude models consistently delivered more accurate answers than human participants when responding to verifiable questions.
At Code With Claude, which also saw the launch of the new Claude Opus 4 and Claude Sonnet 4 models, Amodei reiterated his belief. According to TechCrunch, he responded to a question by saying, “It really depends on how you measure it, but I suspect that AI models probably hallucinate less than humans, but they hallucinate in more surprising ways.”
The upgraded Claude 4 models mark a significant milestone in Anthropic’s push toward artificial general intelligence (AGI), with improvements in memory, code generation, tool use, and writing quality. Claude Sonnet 4, in particular, scored 72.7% on the SWE-Bench benchmark, setting a new bar in software engineering performance for AI systems.
Despite the progress, Amodei was quick to clarify that hallucinations have not been eliminated entirely. In open-ended or loosely structured contexts, AI models are still prone to errors. He stressed that context, prompt phrasing, and use case critically influence a model’s reliability, especially in high-stakes scenarios like legal or medical advice.
His comments come in the wake of a courtroom incident where Anthropic’s Claude chatbot produced a false citation in a legal filing during a lawsuit involving music publishers. The company's legal team later had to apologise for the mistake, underscoring the lingering challenges around factual consistency.
Amodei also emphasised the need for clearer metrics across the industry. With no standard definition or benchmark for what constitutes a hallucination, measuring and ultimately reducing these errors remains difficult. “You can’t fix what you don’t measure precisely,” he warned.
While AI models are making strides in factual accuracy, Amodei’s remarks serve as a reminder that both human and machine intelligence have their flaws, and that understanding, measuring, and mitigating those flaws is the next frontier in AI development.