Apple Research: GenAI Cannot Think Like a Human
- By Paul Mah
- June 11, 2025

GenAI can write well and appears to generate well-reasoned responses. But under the surface, even the most sophisticated models today lack cognitive abilities. What’s more, they eventually experience a complete “accuracy collapse” when tasked with higher-complexity problems.
This was the conclusion of a study by researchers at Apple released over the weekend. Titled “The illusion of thinking,” the paper detailed Apple’s efforts as it tested AI models such as Claude 3.7 Sonnet and DeepSeek-R1 in controlled environments.
Breaking down the study
The research involved the Apple team creating a series of puzzle environments with varying levels of complexity. Both traditional large language models (LLMs) and more advanced large reasoning models (LRMs) were used to attempt to solve these puzzles.
The study found that standard LLMs were more accurate and delivered better results with fewer compute resources for simple tasks compared to LRMs. When complexity increased to a moderate level, models equipped with reasoning capabilities gained the advantage and outperformed the LLMs. However, both types of models failed completely as complex
The researchers argue that state-of-the-art LRMs today fail to generalize problem-solving capabilities, with accuracy collapsing to zero beyond certain levels of complexity. This suggests that the current practice of evaluating LRMs using established math benchmarks is questionable, and that their propensity for accuracy collapse could render outputs unreliable.
The researchers wrote: “We find that there exists a scaling limit in the LRMs’ reasoning effort with respect to problem complexity, evidenced by the counterintuitive decreasing trend in the thinking tokens after a complexity point.”
Enough for massive change
However, prominent AI thought leader Ethan Mollick felt that the conclusions from the Apple research were “vastly overstated.” While he does not dispute the findings, Mollick wrote in a LinkedIn post that AI is not going away, and that existing AI models are good enough to drive massive changes even if the technology stopped advancing today.
In a separate post on the same day, he also observed that Apple isn’t showing signs of catching up in the AI race yet. He wrote: “Apple doesn't report benchmarks for their AIs… But even by their standards, Apple's latest on-device models are mostly worse than the open Gemma 3-4B from Google or Qwen 3-4B, which are also AIs that can run on a phone.”
It’s worth noting that we didn’t have LRMs two years ago. It’s entirely possible that the accuracy collapse observed by Apple reflects the current limits of LRMs. It remains to be seen whether new techniques can bring GenAI significantly further along.
The research paper can be accessed here (pdf).
Image credit: iStock/doble-d
Paul Mah
Paul Mah is the editor of DSAITrends, where he report on the latest developments in data science and AI. A former system administrator, programmer, and IT lecturer, he enjoys writing both code and prose.