Version history
v2.0.6 · 2026-07-30 · Our social preview card was still showing pre-v2 numbers, weeks after the numbers changed. Anyone sharing a link to this site saw a card reading 14 live benchmarks, 20 models tracked and “multimodal”. The truth is 10, 21 and Visual Reasoning. The image file on our server was in fact correct and had been regenerated with the rest of the v2 release; what was wrong was our assumption that replacing a file replaces what the world sees. Social platforms cache preview images against the URL, so re-rendering the same filename changes nothing for anyone who has already shared the link, and we had verified the file rather than the card. Every preview image now lives at a versioned URL, the homepage card and all 21 model cards, which forces every platform to fetch it again. The previous files stay where they are, so an existing embed shows an old image rather than a broken one. The card generator carries the version token now, so this is a one-line change next time rather than a rediscovery. This is the same failure as the human-ceiling counts two releases ago: a rendered artifact that nobody re-derived after the thing underneath it changed. The site text was clean, and we checked: no page description, title or preview text anywhere on the site still carries the old figures.
v2.0.5 · 2026-07-29 · Tapping a specialty tab on a phone gave you a wall of text instead of scores. The specialty-index notice rendered 586px tall on a 375px screen, 72% of the viewport, and pushed the leaderboard 742px below the tab bar. A reader tapping Coding to see coding scores got a screenful of methodology and had to scroll to find a single number. Below tablet width these notices now collapse to a summary with a More expander, and the full text is one tap away in place. Collapsed, the notice is 151px and the table sits 307px from the tabs, so it is reachable in one short scroll. Nothing was deleted and nothing was softened. The visible summary is a compression, not a gentler version: it still says the specialty index is a 0 to 10 scale rather than the AGI Score, that 10 means a perfect score on every benchmark in the set, and that this is not a human-parity claim. Those are the three things a reader has to know before reading the number, so they stay visible whether or not anybody taps More. The same treatment went to the Value view note about cost being an estimate rather than a measured bill, which kept its estimate caveat in the collapsed line. Switching tabs always re-collapses, so no tab inherits the previous one's expanded state. The expander is a real button with aria-expanded, works without hover, and the Back to AGI Score control stays visible while collapsed. Desktop is unchanged: it has the room, so it shows the full text and no expander at all.
v2.0.4 · 2026-07-29 · The mobile tap states would not have rendered on an iPhone. The new section pills carry a pressed state, and it verified correctly in desktop Chromium, which is exactly the wrong place to check it. iOS Safari declines to apply the CSS active state to a link unless a touch listener exists on the element or one of its ancestors, so on the only kind of device that has a touch screen the pills would have looked inert when tapped. One empty listener on the page body fixes it for every tap target on the site. Recorded because verifying a touch behaviour in a mouse browser and calling it done is the kind of shortcut that ships a defect, and because the same trap applies to anything interactive we add from here.
v2.0.3 · 2026-07-29 · The site had no mobile navigation at all. On a phone the header rendered a wordmark and a Contribute button and nothing else. The six section links were desktop-only and the version pill was hidden below tablet width, so a visitor on a phone had no route to Corrections, Methodology or anything else, on a page that is otherwise one continuous scroll, and could not see which release they were reading. Small screens now get a compact sticky bar with the version pill visible and a scrolling row of section pills beneath it, carrying the same six anchors as the desktop nav. A row rather than a hamburger: there is no open state to get stuck, nothing to trap keyboard focus, and every section stays one tap away instead of two. Phones also get a back-to-top button once you are far enough down that scrolling back is a chore. Section anchors were landing behind the header at every width, because nothing on the page set a scroll margin; jumping to Corrections put its own heading underneath the sticky bar. Fixed for both layouts. Nothing else moved: the desktop bar is unchanged, and no score, weight or benchmark is touched by this release.
v2.0.2 · 2026-07-29 · We were overstating how much of the scale is anchored to humans. Three places on this site claimed that three or four scored benchmarks carry a measured human ceiling, and named OSWorld, FrontierMath and SimpleBench among them. The correct number is one. Of the ten benchmarks on the board, only GPQA Diamond (0.81) carries a measured human ceiling; the other nine are scored against the benchmark maximum, where 100 means a perfect score and no human claim is made at all. The other three did pass the same ceiling audit, and none of them is on the board: OSWorld (0.72) was retired in this very release, and FrontierMath (0.35) and SimpleBench (0.837) have no harvested coverage yet. The claim was wrong in the AGI definition modal, in the methodology summary and in the full methodology, and the three did not even agree with one another. This is the same class of error as the correction we published two versions ago: a number that survived because nobody re-derived it after the thing underneath it changed. The AGI Score is a mixed scale, and it is a good deal more mixed than we were saying. Also in this release: the DeepSeek apology is signed by Barak Laniado, founder and CEO, and its corrections contact is now an email address you can actually write to rather than a link back into the site. The calibration constants box no longer reads “Current as of v1.4” under a v2 banner; the constants are unchanged and were re-confirmed for v2.0.0, which is what it now says. Three hover styles on the corrections card and twelve layout classes on the apology page were also missing from the compiled stylesheet, which is purged to what the leaderboard uses, so they had been failing silently.
v2.0.1 · 2026-07-29 · The DeepSeek apology gets its own page. The correction we published in v2.0.0, an ARC-AGI-2 score attributed to DeepSeek V4 Pro that carried no source at all, now has a dedicated page at One count in that entry was also wrong. It said three further cells were voided in the same pass and then referred to four of them in the next sentence. Four further cells were voided, five in total including the DeepSeek one. Corrected here rather than quietly. · 2026-07-29 ·The correction we published in v2.0.0, an ARC-AGI-2 score attributed to DeepSeek V4 Pro that carried no source at all, now has a dedicated page at /corrections/deepseek-arc-agi-2 , linked from its entry in the log above. An apology buried as one card among several is easy to walk past, and this one should be readable, citable and linkable on its own. The page sets out what we published, why it was wrong, what it did and did not affect, and what changed in the process so that it cannot recur: under methodology v2 a recorded source is a condition of scoring rather than an expectation, and all 156 scored cells carry one. The wording of the apology itself is unchanged and identical in both places.It said three further cells were voided in the same pass and then referred to four of them in the next sentence. Four further cells were voided, five in total including the DeepSeek one. Corrected here rather than quietly.
v2.0.0 · 2026-07-29 · Methodology v2. Every score falls, and no model got worse. Two changes drive it. First, human ceilings. Because a score is normalised as raw divided by a human ceiling, that ceiling is not only an anchor for what counts as human level, it is a volume knob: it multiplies both the level and the spread of a benchmark, and so how hard that benchmark pushes on its component. Most of ours had no measurement behind them. Humanity’s Last Exam sat at 0.50, a number our own internal audit described as a policy floor rather than a finding, and that setting quietly made HLE count double. From now on a ceiling is either a published human result under a protocol comparable to the models’, or it is simply the benchmark maximum and we make no human claim at all. Only four survived the test: GPQA Diamond at 0.81, OSWorld at 0.72 (which we raised, having carried 0.85 against our own recorded measurement), FrontierMath at 0.35, and SimpleBench at 0.837, where we had rounded a human baseline up. Eleven moved to 1.00. Three component scores that previously sat above 100, on a scale whose 100 is meant to be the human mark, no longer do. Second, Agency was rebuilt on four legs across three independent evaluators: SWE-bench Verified, Terminal-Bench 2.1, τ³-Banking and LiveBench Agentic Coding. Terminal-Bench now comes from vals.ai rather than Artificial Analysis, cutting our reliance on any single evaluator. τ³-Banking enters at benchmark maximum, so its true range shows: the best model completes about a third of these stateful banking workflows. Retired: SWE-bench Pro, OSWorld, BrowseComp, Tau-bench retail and airline, Aider Polyglot, Terminal-Bench 2.0. MCP Atlas demoted to informational. The Tool Use tab is withdrawn until a second clean non-coding benchmark exists. Scores fell by about 6 points from the ceiling change and further from the Agency rebuild; ordering shifted only locally, and the top two are unchanged. Third, evaluator concentration. Artificial Analysis had been supplying 45% of the whole board, 97% of Knowledge and 100% of Visual Reasoning, so one evaluator changing terms would have taken two components to zero. GPQA Diamond and MMMU-Pro moved to vals.ai alongside Terminal-Bench 2.1, taking Artificial Analysis to 26% of the board and splitting Knowledge 52/48 between two evaluators. Every switch was measured before it was made rather than after, and each offset is published rather than absorbed: GPQA Diamond −0.30pp with 1.38pp scatter across all 20 models, MMMU-Pro +5.71pp with 1.67pp scatter and vals higher on all eleven comparable models, Terminal-Bench 2.1 −7.77pp with 4.54pp scatter and vals lower on 18 of 19. One row is held rather than used: vals reports Grok 4.5 at 61.8 on MMMU-Pro, below Grok 4.3 at 83.1 and below xAI’s own non-reasoning variants, which is not a capability measurement we can account for, so we do not use it. Grok 4.5 consequently has no Visual Reasoning cell and falls three places. Visual Reasoning is still single-source, having moved from wholly Artificial Analysis to wholly vals.ai: it rests on one benchmark, so this relocates the dependency rather than removing it, and we would rather say so than imply it is solved. Also in this release: the Reasoning specialty tab is withdrawn. It ran on ARC-AGI-2 and AIME 2025, but AIME 2025 is scored for a single model and sits at 98-99% where it appears, so the tab was ranking on ARC-AGI-2 alone while labelled as using two benchmarks. Withdrawn under the same rule as Tool Use rather than kept with a misleading label; it returns when AIME 2026 is harvested. Cursor Composer 2.5 is removed from the board: it is a product system rather than a model, so under v2 it scores no cells at all, and listing it implied we were waiting on evidence that was never coming. Specialty tabs now report a 0 to 10 index rather than a 0 to 100 score. On the AGI Score, 100 means the genesis of AGI; a specialty score never meant that, and showing both on the same scale implied they were the same kind of claim. Ten now means a perfect score on every benchmark in that set, with no human comparison implied, because most of those benchmarks have no published human study to compare against. The Multimodal component is renamed Visual Reasoning. It rests on a single scored benchmark, MMMU-Pro, which is an exam built on diagrams and charts, so calling 11% of the score “multimodal” claimed a breadth of perception we do not measure. Visual Reasoning says what it is: interpreting image data, which underpins reading a chart, extracting a figure from a document, or working from a screenshot. The AGI definition was revised in the same pass and for the same reason. It previously excluded “sensory perceptions whatsoever”, which contradicted our own scoring of a diagram-based benchmark. Rather than bolt vision on, which would have raised the obvious question of why sight and not hearing or smell, we removed the sensory clause and let the physical-body clause carry the weight: what reaches a model as data is in scope, what needs a body to acquire is not. The definition got shorter. The corrections summary counters have also been replaced. They had been hardcoded since May and one of them read “1 model entered scoring” for two months without meaning anything; they are now derived from the scoring ledger on every release.
v1.11.28 · 2026-07-25 · Corrections to our corrections. We audited this log against our own commit history and found that two entries were wrong. Both times we compared two different configurations of the same benchmark and reported the difference as a laboratory overstating its results. DeepSeek did not overstate GPQA Diamond: the 72.9% we published as an independent contradiction is DeepSeek’s own non-thinking-mode figure, reaching us through a secondary aggregator and set against their thinking-mode result. Measured like for like, their self-report sits 1.3 points above independent measurement rather than 17.2 points above. Moonshot did not overstate Humanity’s Last Exam: their 54.0% was explicitly labelled as a with-tools result, and the same post published 36.4% without tools, within half a point of independent measurement. Our source policy had excluded them on the strength of that misreading. Both entries are retracted and both originals are retained rather than deleted. We have also withdrawn a laboratory promotion test that the policy described but that was never run, and corrected an entry which presented source relabelling as new independent measurement. No score, weight, human ceiling or source tier changed in this release.
v1.11.27 · 2026-07-25 · Claude Opus 5 and Anthropic roster correction. Added Claude Opus 5 on five independent, variant-distinct frozen-v1 cells: ARC-AGI-2 90.4%, SWE-bench Verified 97.0%, HLE 52.59%, GPQA Diamond 93.23%, and MMMU-Pro 84.74%. The unchanged eligibility engine places it Provisional at 97.31 because one Agency cell and no accepted LiveBench row leave two thin components. The available LiveBench EAP row remains held pending public-model identity resolution; Terminal-Bench 2.1, ARC-AGI-1 and ARC-AGI-3 remain outside frozen-v1 scoring. Anthropic’s active latest-two roster is now Opus 5 and Opus 4.8. Fable 5 leaves the current roster under an identity-integrity exception because fallback-enabled, fallback-disabled and generic rows cannot be combined into one standalone model; Opus 4.7 thinking and base retire as older generations. All historical evidence remains preserved, with no transfer or imputation. No formula, weight or source tier changed.
... continue reading