Segmenting Robot Video into Actionable Subtasks
(news.ycombinator.com)
121.
122.
Ornith-1.0: self-improving open-source models for agentic coding
(news.ycombinator.com)
123.
GLM 5.2 beats Claude in our benchmarks
(news.ycombinator.com)
124.
Does the Nvidia App Hurt Performance? We Benchmarked It.
(techspot.com)
125.
Does the Nvidia App Hurt Performance? We Benchmarked It.
(techspot.com)
126.
127.
Lies, Damn Lies and Database Benchmarks
(news.ycombinator.com)
128.
An Introduction to YOLO26
(news.ycombinator.com)
129.
I gamed on 2026’s best Snapdragon and Exynos flagship phones — and the benchmarks lied
(androidauthority.com)
130.
131.
PostgresBench: A Reproducible Benchmark for Postgres Services
(news.ycombinator.com)
132.
The best 3D scanners 2026 — the top performing models we've benchmarked
(tomshardware.com)
133.
GLM 5.2 Performance Benchmarks
(news.ycombinator.com)
134.
135.
136.
Google’s Android coding tests reveal an unexpected Gemini 3.5 Flash weakness
(androidauthority.com)
137.
Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
(news.ycombinator.com)
138.
139.
MTG Bench: Testing how well LLMs can play Magic
(news.ycombinator.com)
140.
What AI benchmarks miss about real-world performance
(venturebeat.com)
141.
142.
143.
144.
145.
Benchmarks in Leipzig
(news.ycombinator.com)
146.
147.
148.
Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities
(news.ycombinator.com)
149.
These LLMs are the best at resisting Russian propaganda
(arstechnica.com)
150.
Today's top topics:
polymarket