A new Epoch AI report by Emberson and Roodman finds that the price of achieving a given level of AI performance has dropped roughly 47% per quarter over the past three years, a 13-fold annual decline. The report cites OpenAI's o3, which cost about $0.30 per question to score 75% on GPQA Diamond in January 2025, compared with a much newer model reaching similar scores for a fraction of a cent by mid-2026.
marginalrevolution.com
· 2026-09-23
The developer behind the iMessage AI assistant Olly, which has processed over 18 million messages with roughly a third routed through OpenRouter's open-source models, published a technical breakdown of problems encountered at scale. The core issue: OpenRouter can send identical model requests to any of about 20 different hosting providers, each running the same weights but with different precision, optimizations and tool parsers, producing measurably different benchmark results. For DeepSeek V4 Flash, GPQA Diamond and TAU-Bench Airline scores varied by several points across providers on the same day.
mmoustafa.com
· 2026-09-09
Artificial Analysis released an interim update to its AI model benchmark, adding a new private agentic knowledge-work test called AA-Briefcase and a long-context PDF reasoning test built by Surge covering 4,592 pages. The update also drops GPQA Diamond, which the company says is now saturated, and adds more held-out test sets and stronger grading infrastructure to make gaming the benchmark harder. The change comes eight months after Index v4 launched, as the team continues building toward a fuller v5 release.
artificialanalysis.ai
· 2026-09-05