When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
(news.ycombinator.com)
1.
2.
Homebench – Benchmark local LLMs for speed, memory, and quality
(news.ycombinator.com)
3.
Orca-Bench: How Ready Are Language Model Agents for Oncall?
(news.ycombinator.com)
4.
5.
DeepSeek-V4-Flash Update
(news.ycombinator.com)
6.
Register deprivation: spills and runtime under forced register scarcity
(news.ycombinator.com)
7.
Does Speaking to Agents Like Cavemen Save 65% of Tokens? We Test
(news.ycombinator.com)
8.
Benchmarking Opus 5 on SlopCodeBench
(news.ycombinator.com)
9.
Show HN: Optimize and serve models with Fable quality at half the cost
(news.ycombinator.com)
10.
UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
(news.ycombinator.com)
12.
13.
Geekbench 7
(news.ycombinator.com)
14.
15.
16.
Samsung’s latest foldables are having a benchmark controversy, right out of the gate
(androidauthority.com)
18.
Any text-to-SQL benchmark should address difficulties of real-world data stores
(news.ycombinator.com)
19.
The Last MPEG-4 Visual Patent Has Expired
(news.ycombinator.com)
20.
21.
22.
Kimi K3, and what we can still learn from the pelican benchmark
(news.ycombinator.com)
23.
Bun vs. Deno vs. Node.js 2026: Real Benchmarks Mislead
(news.ycombinator.com)
24.
Assassin's Creed Black Flag Resynced: 50 GPU Benchmark
(techspot.com)
25.
Assassin's Creed Black Flag Resynced: 50 GPU Benchmark
(techspot.com)
26.
GLM 5.2 is nearly as accurate as a human book keeper
(news.ycombinator.com)
27.
Benchmarking coding agents on Databricks' multi-million line codebase
(news.ycombinator.com)
28.
29.
30.
Google updates Android Bench with new LLMs, but Gemini still lags behind
(arstechnica.com)