Coherence Without Comprehension
LLMs produce fluent text without grounding. What hallucination benchmarks reveal about the gap between sounding coherent and actually understanding.
Introduction

In Foucault’s Pendulum (1988), Umberto Eco constructs a compelling narrative in which three editors from Milan — Casaubon, Belbo, and Diotallevi — become deeply entangled in a conspiracy theory of their own invention named “The Plan.” Central trigger of this descent is Abulafia, a computer they use to generate random connections between disparate historical texts. Initially a playful tool, Abulafia’s output becomes a pseudo-oracle, spitting out increasingly elaborate associations between the Templars, Rosicrucians, and the occult. Though the editors know the machine merely recombines texts, they eventually begin to believe the texts it generates. “The Plan” conspiracy gradually spirals into obsession, drawing real-world consequences, including paranoia, delusion, and ultimately, murder — when Belbo is killed by fanatics who have come to believe in the very conspiracy the editors created.
Written in 1988 Abulafia was of course not a direct reference to today’s “intelligent machines”. The machine was described as rather deterministic and it simply combined the fragments of human-authored text into something that appeared coherent. Nevertheless, in its combinatorial mimicry of knowledge, Abulafia eerily prefigures the behavior of modern large language Models. These Models are trained on socially constructed language about the world. They produce what appears to be knowledge by remixing what humans have said, thought, and imagined. The results often dazzle with plausibility, but the boundary between signal and noise is very thin. LLMs hallucinate — fabricate plausible-sounding but false information. Like Abulafia, they are engines of synthesis without referential grounding.
Coherence Without Comprehension
In the recent years there has been a large body of empirical research to systematically quantify how often large language models hallucinate, and the results are quite intriguing. For example, in DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation authors, introduced a benchmark of over 75,000 question-answer pairs spanning eight factual domains — ranging from Nobel Prize winners and census data to university rankings and math. The dataset was explicitly designed to elicit definitive answers to concrete questions involving dates, names, locations, or numeric values, with each question paraphrased 15 times to test for response consistency. 3 metrics were measured :
- Factual Accuracy (Does the model provide the correct answer?)
- Prompt misalignment (Does the model follow the instruction?)
- Response Consistency (Does it answer the same question reliably across paraphrases?)

Factual hallucination rates (FCH rate in the charts) ranged from 59% to 82% across models, with particularly severe failure modes in domains requiring numerical precision (e.g., census statistics, math problems, and university rankings). Prompt misalignment — where a model strayed from the question’s format or intent — occurred in 6% to 95% of outputs, depending on the model. Even consistency was lacking: models often gave contradictory answers to paraphrased versions of the same question, with consistency scores ranging from 21% to 63%. Notably, high-profile models performed better than open-source alternatives but still exhibited non-trivial hallucination rates, especially when forced to recall precise facts.
These findings underscore a critical point: LLMs do not merely occasionally fabricate information — they do so consistently at rates that, in many contexts, would be completely unacceptable for institutional knowledge systems.
Naturally, what makes this more than a technical problem is that human knowledge systems themselves are not neutral or stable. In The Archaeology of Knowledge Foucault argued that what counts as knowledge is shaped not by timeless facts but by historical conditions and institutional power structures. There is no clean division between language and belief, between discourse and truth. Today, LLMs are trained on the sedimented layers of human discourse: journalism, blogs, scientific papers, Reddit threads, fan fiction. These are not mere traces of reality but realities of their own often shaped by ideology and bias. When a model generates a text, it is not generating a neutral representation — it is sampling from a contested, constructed archive of messy human discourse.
Confusing the discourse with the knowledge and understanding is precisely the epistemological trap that modern LLM creators fall into. For example, Ilya Sutskever, co-founder of OpenAI, highlighted in an interview:
“On the surface, it may look like we are just learning statistical correlations in text. But it turns out that to compress them really well, what the neural network learns is some representation of the process that produced the text. This text is actually a projection of the world; there is a world out there and it has a projection on this text.”
This view suggests that effective compression — the ability to predict sequences well — requires models to internalize something akin to a latent world model. Basically implying that neural networks, by learning from vast amounts of language, implicitly reconstruct the causal or generative processes underlying human experience.
However, the problematic nature of working with such projections of the world was already beautifully illustrated over 2,000 years ago in Plato’s allegory of the cave. In the allegory, Plato describes prisoners who have spent their entire lives chained in darkness, forced to watch shadows cast on a wall by objects they cannot see.

To them, these flickering shadows are not mere projections — they are reality itself. They speak of them, reason about them, even build theories around them, all without ever gaining knowledge about the real forms that cast them. Large language models are prisoners of the cave par excellence and there is a growing body of empirical evidence to prove it.
In their 2025 study “What Has a Foundation Model Found?”, Vafa et al. (Harvard & MIT) challenge precisely this premise. They ask a direct question: does good predictive performance imply the acquisition of an underlying world model? Using a methodology they call inductive bias probing, the authors evaluate whether foundation models trained on structured data — such as orbital trajectories governed by Newtonian physics — internalize the correct physical laws.
And it turns out that even when models accurately predict planetary orbits, they fail to exhibit any inductive bias toward Newtonian mechanics. Instead, they learn task-specific heuristics that mimic the data but do not generalize. When prompted to extrapolate on related tasks, these models hallucinate inconsistent “force laws,” revealing no coherent internalization of gravitational principles. In other domains — such as lattice problems or games — similar patterns emerge: models latch onto surface-level legal moves or token regularities rather than deeper structures.
As Vafa et al. put it, foundation models can “excel at their training tasks yet fail to develop inductive biases towards the underlying world model.” This undermines the notion that sequence modeling alone — even at scale — is sufficient for capturing latent, causal structures in the world. Akin to Plato’s cave allegory, Large Language Models are aware only of the shadows — the linguistic traces humans have left behind. Their understanding, is of surfaces and silhouettes: human experiences, flattened into tokens.
Why should we care?
Unlike Abulafia — which generated random associations, had no agency of its own, and no actuation point in the physical world — large language models today are ubiquitous. They are proprietary, privately owned and embedded in customer service chatbots, legal tools, creative platforms, educational apps, and decision-making systems in hiring, healthcare, and governance.
These are millions of hallucinating Abulafias distributed across the globe, each generating persuasive language, each capable of nudging belief or shaping action. Their outputs are often treated not as probabilistic guesses, but as authoritative statements, while the users often have little or no understanding of how these models work or what assumptions underlie their training. Hence, the experiment is no longer confined to linguistic probabilities in a lab; these systems operate at scale, act in real-time, and influence real decisions. To undermine this potential danger, it is often argued that LLMs lack physical modes of actuation — and therefore lack real capacity to cause harm in the physical world. However, the danger is far from hypothetical. Chatbots have contributed to suicides, disinformation cascades, misdiagnosed mental health crises, and prompted dangerous behavior. LLMs don’t need a body to affect the physical world, they can borrow yours.
If a book can shape collective consciousness, inspire religions, rewire neural pathways, and shift belief systems — then so can a text in a chatbot. After all, both are just language on a surface.
There is of course a crucial difference: human written text is part of discourse; chatbots are immune to it. A text in a book or a paper enters the public domain of critique and scrutiny. We can all read the exact same piece, argue over its meaning, interpret it in different contexts, and refine our understanding through shared deliberation. But with LLMs, there is no stable artifact. My chatbot and your chatbot are not the same. Each prompt generates a private, non-repeatable instance. There is no canonical base to analyze or a shared object to dispute. This creates a rupture. We are outsourcing belief-formation to these systems while we cannot interrogate the model’s priors, trace its citations, or argue with its intent. It is manufactured anew in every interaction. And yet, while each interaction is private and fragmented, the system behind them is anything but. These models are not decentralized independent instances whispering into the ears of billions. They are centralized systems, built, owned, and operated by a handful of corporate actors, each controlling the system prompts and outputs. What appears to be a personal tool is, in reality, a globally scaled instrument — one capable of producing the same belief-framing narrative across vast populations, quietly aligning thought and behavior at scale.
A chatbot just like book cannot act on its own, but both can persuade and human actions will follow. The LLM does not need a physical body — its body is distributed across its user base made of the humans who act on its outputs. LLMs don’t need bodies, when they can animate yours through clicks, shares, investments, hires, votes, and self-diagnosis. A very gentle nudge in a model’s system prompt or training data can cause millions of purchases, alter facts and trigger revolutions. And like the editors in Foucault’s Pendulum, users are increasingly seduced by patterns that seem increasingly more meaningful.
Conclusion
In the Eco’s novel, Abulafia was not dangerous because it fabricated texts — it was dangerous because it persuaded. Its coherence was mistaken for truth and fluency for insight. The same can be said of today’s LLMs. They do not uncover reality; they reshuffle the semiotics with misplaced confidence. Their persuasiveness is not proof of understanding; it is the result of optimization and reinforcement training with humans in the loop.
In the novel, the game turns deadly for the authors of the conspiracy in our world, the stakes are even higher. The Abulafias are no longer tucked away in a dusty office — they live in our phones, our workflows, our institutions affecting decisions. We are all Belbo now, whispering questions into the machine yet we don’t know who owns “The Plan” — or whether the Plan is already shaping us.
Much of these arguments may come across as doom and gloom, however, I do believe that large language models are promising tools — potentially transformative in a number of fields. While all innovation carries risks, I am convinced that the current trajectory demands more scrutiny. What is concerning is not necessarily the technology itself, but the structure of its development. Its widespread accessibility paired with the immense resource intensity required to train cutting-edge models — both in compute and talent. A handful of companies control the knowledge, the infrastructure, the models, and increasingly, the discourse around AI itself. This centralization concentrates the intellectual direction of the field. Too few brains, in too few rooms, are making decisions that have the potential to affect too many.
References
Allegory of the cave - Wikipedia
https://en.wikipedia.org/wiki/Foucault%27s_Pendulum