Skip to content
Tech News
← Back to articles

Anthropic: Introducing The Conceptual Reasoning Index

read original more articles
Why This Matters

The introduction of the Conceptual Reasoning Index (CRI) marks a significant advancement in evaluating AI's ability to handle complex reasoning tasks essential for AI safety and risk mitigation. This development enables researchers and industry leaders to better assess and improve AI models' capabilities in areas critical for managing future AI risks, ultimately supporting safer and more reliable AI deployment.

Key Takeaways

Emery Cooper1, Caspar Oesterheld1, Chi Nguyen1, Alex Kastner1,

Joe Benton2, Ethan Perez2 August 12, 2026 1Redwood Research; 2Anthropic

tl;dr A core hope for managing AI risks is that AIs will help us understand our situation, plan for what lies ahead, and develop risk mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning benchmarks. You can request access to our primary conceptual dataset, LMCA, through this form. We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai, where you can also find more details on our methodology. We will keep the website up to date as both new models and benchmarks are released.

This work was done in collaboration with Anthropic.

Background

Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities. One way to influence this might be to selectively improve models' relevant skills, such as reasoning about how to govern and align AI and how to avoid catastrophic cooperation failures involving AI.

Current AI training depends heavily on abundant data and reliable feedback on the model's performance. Models are therefore typically worse at tasks that cannot be empirically or mathematically verified.1 Unfortunately, reducing risks from advanced AI involves many such tasks:

Much AI safety work involves reasoning about AIs more generally capable than any human. There's no obvious reference class for this and no clear way to model it.

We might have to get some things right the first time. For example, if a mistake leads to AGI takeover or an AI-assisted coup, we might not find out until it's too late. Similarly, many decisions (e.g., which research agendas to prioritize, which governance interventions to pursue) play out over long timescales, such that empirical feedback might not arrive early enough to help.

Lastly, some important questions, such as which values AIs should have, may lack a ground truth entirely (yet we still think progress can be made by arguing about these questions).

... continue reading