Skip to content
Tech News
← Back to articles

Is AI reasoning right for the wrong reasons?

read original more articles
Why This Matters

The rapid advancements in AI reasoning models, especially LRMs, highlight both their impressive capabilities and the ongoing debate about whether they truly 'think' or simply mimic reasoning. This has significant implications for the tech industry and consumers, as it influences trust, application scope, and future development of AI systems. Understanding these nuances is crucial for responsible deployment and innovation in AI technologies.

Key Takeaways

I’ll just say it: What the hell is going on with AI “reasoning”?

Sorry for the air quotes. That punctuational side-eye was more common in 2024, when the specially trained cousins of LLMs now known as “large reasoning models,” or LRMs, were still new. Nowadays it may seem downright churlish, though, given that a “general-purpose reasoning model” from OpenAI solved a famous open mathematical research problem in one shot in May 2026. Still, I’m not sure how else to acknowledge my intellectual whiplash over the scientific interpretation of what these AI systems are actually doing.

Reasoning comes in many technically defined forms, but the basic procedure is easily recognizable: arriving at a sound conclusion by linking together intermediate steps that logically follow from each other. We do this with thoughts; LRMs use so-called chains of thought, a term of art for the streams of synthetic text that the models emit before arriving at an answer to a complex query. One minute, the idea that AI could reason via these chains was being prominently and credibly critiqued (by a team of researchers from Apple) as an “Illusion of Thinking” subject to “complete accuracy collapse” under surprisingly simple conditions. The next minute, LRMs were bagging gold medals at the International Mathematical Olympiad, a feat so challenging that “even very successful mathematicians and scientists may well highlight [it] on their CVs all their lives,” as the scientist and AI critic Gary Marcus and Ernest Davis wrote in 2025. If that’s not a sign of “real” reasoning, what is?

In philosophy, “qualia” refers to the subjective qualities of our experience: what it’s like for Alice to see blue or for Bob to feel delighted. Qualia are “the ways things seem to us,” as the late philosopher Daniel Dennett put it. In these essays, our columnists follow their curiosity, and explore important but not necessarily answerable scientific questions.

But wait — soon after, more research, from the Santa Fe Institute, showed that LRMs can crush even carefully designed benchmarks for reasoning (like a collection of analogy-like visual puzzles) using mere “surface-level ‘shortcuts.’” What they were doing looked less like generalizable reasoning than just gaming the system. Then, as if on cue, another “hold my beer” moment: Google DeepMind and the mathematician Terence Tao (the GOAT!) used AI to rediscover or improve the solutions to 67 problems “spanning mathematical analysis, combinatorics, geometry, and number theory.” Deal with it, haters!

What about additional evidence that LRMs can’t reason reliably, even when they possess the necessary algorithm and computational budget to do so, and suffer from a list of scientifically documented failure states long enough to use as a Slip ’N Slide? Whatever — I guess that’s just “jagged intelligence” for you (AI-speak for “when it works, it works”).

And so it went from late 2025 into 2026. I’ve been a science journalist for 20 years and an AI journalist for half of that, so I know better than to expect tidy consistency out of rapidly advancing research. But even for me, this back-and-forth has been a bit much. To quote Al Pacino in The Insider, “I’m getting two things: pissed off, and curious.” I don’t believe there’s fraud to be found here. I just want to know which way is up. Can AI reasoning somehow be both BS and not at the same time? And if so, how on Earth does that work?

I knew just who to call first.

Melanie Mitchell’s career in AI stretches back to the 1980s, but lately she’s earned a reputation as an au courant AI truth teller, penning lucid explainers for Science and her widely read newsletter, as well as conducting research at the Santa Fe Institute. (The study about “surface-level ‘shortcuts’” is hers.) When I asked her what we actually know about AI reasoning, her answer was brief enough to fit on an index card.

“Number one: It works. It improves things,” she said, referring to LRMs’ superior accuracy on reasoning tasks compared to LLMs. “Number two: The actual text that’s generated” — i.e., the chain of thought that every LRM is trained to produce to improve its performance — “isn’t necessarily faithful to what’s going on [inside the model]. And number three: A lot of that text isn’t even useful. You can actually take it out.”

... continue reading