Skip to content
Tech News
← Back to articles

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

read original more articles
Why This Matters

Terminal-Bench-Science represents a significant step forward in evaluating AI's ability to assist with complex scientific research workflows. By setting a high standard based on real scientific tasks, it encourages the development of AI tools that can genuinely augment researchers' capabilities, potentially accelerating scientific discovery and innovation. This benchmark underscores the importance of aligning AI development with actual scientific needs, benefiting both the industry and consumers by fostering more effective and intelligent research assistants.

Key Takeaways

Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI.

Terminal-Bench-Science is a benchmark led by researchers at Stanford University and built by the team behind Terminal-Bench in collaboration with domain experts from a range of scientific disciplines and research institutions around the world. It measures the AI agent capabilities through a diverse set of challenging, expert-curated workflows drawn from scientific research.

Terminal-Bench-Science is a continuous benchmark that evolves alongside frontier AI, creating a feedback loop between scientific needs and AI development. Our first release includes 70 tasks from the life, physical, Earth, mathematical, and engineering sciences. The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1.

Terminal-Bench-Science 0.1 Leaderboard 0% 25% 50% 75% 100% Claude Opus 5 Claude Code 30.0% GPT-5.6 Sol Codex 22.4% Claude Fable 5 Claude Code 21.4% Claude Opus 4.8 Claude Code 10.5% GPT-5.6 Terra Codex 8.6% GLM 5.3 Claude Code 8.1% Kimi K3 Claude Code 7.1% Grok 4.6 Grok Build 7.1% GPT-5.6 Luna Codex 3.3% Resolution rates across 70 scientific workflow tasks on Terminal-Bench-Science 0.1

While Terminal-Bench has driven progress in AI agents for software engineering, Terminal-Bench-Science brings the same ambition to science. Our goal is to drive the development of agents with scientific capabilities that make them useful research assistants. These agents should execute technically demanding and time-consuming workflows, freeing scientists to focus more of their time on the parts of science where human judgment matters most: defining research questions, forming hypotheses, interpreting and validating results, and communicating findings. In this role, AI agents can extend what researchers accomplish and help accelerate scientific discovery.

Achieving this requires benchmarks that reflect real scientific practice, provide verifiable evidence of capability, and evolve alongside the AI frontier.

We need benchmarks drawn from real scientific workflows. Scientific capability should be evaluated on real research practice, not textbook questions or standardized exercises, contributed by practicing scientists themselves. Terminal-Bench-Science gives scientists across domains a direct voice and a shared platform to set the bar for AI progress on the problems they care about. The stakes in science are too high, and its benchmarks must reflect the scientific community's priorities rather than outside interests.

We need verifiable evidence of scientific capability. Without reliable evaluation, we cannot tell whether agent capabilities are improving or where their limitations remain. Terminal-Bench-Science evaluates agents in realistic environments and grades concrete artifacts such as analyses, simulations, proofs, code, and data products with reproducible, task-specific tests.

We need a benchmark that keeps pace with the frontier. Too often, scientific benchmarks are treated as papers to publish rather than mechanisms for driving progress. They are released once and then abandoned as models advance and known limitations persist. Terminal-Bench-Science is a continuous benchmark that evolves alongside the AI frontier. Through regular releases, scientists can contribute new workflows, improve existing tasks, and create a feedback loop between scientific needs and AI development.

Terminal-Bench-Science Feedback Loop SCIENTIFIC COMMUNITY FRONTIER AI AGENTS & MODELS Contribute workflows Evaluate & improve Accelerate discovery Terminal-Bench-Science Feedback Loop

... continue reading