Gemini 3.8 Live and 3.8 Live Extended Thinking
(news.ycombinator.com)
1.
2.
Pion, an agent designed to run any company autonomously
(news.ycombinator.com)
3.
RTK reports token savings, but our cost benchmarks disagree
(news.ycombinator.com)
4.
Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
(news.ycombinator.com)
5.
Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1
(news.ycombinator.com)
6.
So you want to use OpenRouter?
(news.ycombinator.com)
7.
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(news.ycombinator.com)
8.
GPT-6 Astra
(news.ycombinator.com)
9.
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
(news.ycombinator.com)
10.
Ornith-1.5: From Self-Scaffolding to Self-Improvement
(news.ycombinator.com)
11.
12.
Orca-Bench: How Ready Are Language Model Agents for Oncall?
(news.ycombinator.com)
13.
Google updates Android Bench with new LLMs, but Gemini still lags behind
(arstechnica.com)
14.
AI tools can speed up thinking, but evidence still comes from the lab bench
(feeds.nature.com)
15.
Segmenting Robot Video into Actionable Subtasks
(news.ycombinator.com)
16.
Ornith-1.0: self-improving open-source models for agentic coding
(news.ycombinator.com)
17.
18.
19.
20.
Ikea’s newest furniture makes Scandinavian design fun again
(feeds.feedburner.com)
21.
22.
The Disappearance of the Public Bench
(news.ycombinator.com)
23.
SWE-bench Verified no longer measures frontier coding capabilities
(news.ycombinator.com)
24.
Why SWE-bench Verified no longer measures frontier coding capabilities
(news.ycombinator.com)
25.
26.
N-Day-Bench – Can LLMs find real vulnerabilities in real codebases?
(news.ycombinator.com)
27.
Exploiting the most prominent AI agent benchmarks
(news.ycombinator.com)
28.
29.
Today's top topics:
openai
anthropic
apple
ai safety
google
ios 27
dario amodei
iphone 18 pro
nvidia
microsoft