Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
(news.ycombinator.com)
1.
2.
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(news.ycombinator.com)
3.
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
(news.ycombinator.com)
4.
Ornith-1.5: From Self-Scaffolding to Self-Improvement
(news.ycombinator.com)
5.
Today's top topics:
apple
iphone duo
openai
iphone 18 pro
google
anthropic
apple watch series 12
amazon
android
samsung