I was consistently using an xhigh or max effot level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all. And when longer thinking runs did happen, they almost never reached published benchmark levels.
Fable 5 – Median thinking declined in August
Why This Matters
This story highlights a discrepancy between advertised AI model capabilities and their real-world behavior, raising concerns about transparency in how 'effort' or 'thinking' settings actually perform versus benchmark claims. For consumers and developers relying on these settings for complex tasks, this could mean inconsistent or underdelivered performance despite paying for higher-tier compute options.
Key Takeaways
- Median 'thinking' token usage dropped in August despite max effort settings being selected.
- Actual model reasoning depth frequently fell short of published benchmark levels.
- The findings suggest a gap between marketed AI performance and real-world output consistency.
Get alerts for these topics