The Efficiency Trap: Why Statistically Optimal AI Misses Human-Like Understanding
Statistical efficiency and human-like intelligence appear to be fundamentally at odds. New research from CDS and Stanford University reveals that large language models excel at compressing information into tidy categories, but this very efficiency prevents them from capturing the rich, contextual nuances that define human understanding.
The study, “From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning,” was led by Stanford postdoc Chen Shani, and included CDS Research Scientist Ravid Shwartz-Ziv, Stanford professor Dan Jurafsky, and CDS founding director Yann LeCun. The team used an information-theoretic framework to compare how humans and LLMs organize concepts like “fruit” or “bird,” drawing on classic cognitive science datasets that map how people categorize the world around them.
“We wanted to examine how LLMs represent concepts,” Shwartz-Ziv said. “For example, what does the concept of ‘fruit’ mean? You have apples, bananas, and other items within that category. We looked at studies from cognitive science that asked people about the conceptual distance between apple and fruit, and the distance between cucumber and fruit.”
The researchers tested dozens of language models, from small 500-million parameter systems to massive 72-billion parameter models. They found that LLMs could successfully form broad conceptual categories that matched human judgment; i.e., both humans and machines grouped apples with bananas under “fruit.” But when the team examined the internal structure of these categories, a striking difference emerged.
Humans perceive some items as more typical than others within a category. A robin feels more “bird-like” than a penguin, even though both are birds. LLMs struggled to capture these typicality gradients, treating category members more uniformly than humans do.
The most revealing finding came from applying information theory to measure the trade-off between compression and meaning preservation. LLMs consistently achieved what the researchers call “statistically optimal” representations — they compressed information efficiently while minimizing distortion. Human conceptual systems, by contrast, appeared “suboptimal” by these same statistical measures.
“LLMs store more efficient representations compared to humans,” Shwartz-Ziv explained. “They compress all the contextual information that isn’t directly related to the core task. Humans store more information in their representations, which makes them less efficient from an information-theoretic perspective, but this additional information can be valuable for generalization and use in other contexts.”
This difference reflects fundamentally different optimization pressures. LLMs, trained on vast text corpora, learn to maximize statistical regularity and minimize redundancy. Human cognition evolved under different constraints, prioritizing adaptive flexibility, causal reasoning, and the ability to generalize across diverse contexts — even if this means storing seemingly “inefficient” representations.
The findings challenge assumptions about the relationship between linguistic competence and conceptual understanding. “Even though LLMs are proficient with language and words, they are fundamentally different from humans,” Shwartz-Ziv said. “We shouldn’t be misled by the fact that they use language and English. We need to understand that they are not thinking like humans.”
LeCun has argued that current language model paradigms may be insufficient for achieving human-level intelligence. This research provides quantitative evidence for that perspective, showing how aggressive compression — while computationally efficient — may limit the rich, contextual representations that support robust reasoning.
The researchers digitized and released the classic cognitive science datasets they used, making them available for future AI research. The work suggests that developing more human-like AI may require moving beyond pure statistical efficiency toward systems that can maintain the kind of adaptive, context-rich representations that characterize human thought.
“Now that we understand that these differences exist,” Shwartz-Ziv said, “we may want to develop specific training algorithms that better align LLMs with human understanding.”
By Stephen Thomas
Have feedback on our content? Help us improve our blog by completing our (super quick) survey.
