Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
121.
Segmenting Robot Video into Actionable Subtasks (news.ycombinator.com)
122.
Ornith-1.0: self-improving open-source models for agentic coding (news.ycombinator.com)
123.
GLM 5.2 beats Claude in our benchmarks (news.ycombinator.com)
124.
Does the Nvidia App Hurt Performance? We Benchmarked It. (techspot.com)
125.
Does the Nvidia App Hurt Performance? We Benchmarked It. (techspot.com)
126.
Alibaba's model never trained as an agent — and improved agent performance across seven benchmarks (venturebeat.com)
127.
Lies, Damn Lies and Database Benchmarks (news.ycombinator.com)
128.
An Introduction to YOLO26 (news.ycombinator.com)
129.
I gamed on 2026’s best Snapdragon and Exynos flagship phones — and the benchmarks lied (androidauthority.com)
130.
OpenAI Announces Benchmarks for AI Life Sciences Research. Its Best Model Failed 63.9% of the Test (slashdot.org)
131.
PostgresBench: A Reproducible Benchmark for Postgres Services (news.ycombinator.com)
132.
The best 3D scanners 2026 — the top performing models we've benchmarked (tomshardware.com)
133.
GLM 5.2 Performance Benchmarks (news.ycombinator.com)
134.
Why Weibo’s tiny VibeThinker-3B has the AI world arguing over benchmarks again (venturebeat.com)
135.
Z.ai’s open-weights GLM-5.2 beats GPT-5.5 on multiple long-horizon coding benchmarks for 1/6th the cost (venturebeat.com)
136.
Google’s Android coding tests reveal an unexpected Gemini 3.5 Flash weakness (androidauthority.com)
137.
Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks (news.ycombinator.com)
138.
Kimi K2.7-Code cuts thinking tokens 30% — but practitioners say the benchmarks don't check out (venturebeat.com)
139.
MTG Bench: Testing how well LLMs can play Magic (news.ycombinator.com)
140.
What AI benchmarks miss about real-world performance (venturebeat.com)
141.
Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark (venturebeat.com)
142.
AMD fires back at Nvidia, claiming 256-core Zen 6 'Venice' CPU beats Vera by 3.3x in rack-level performance — company shares first estimated EPYC Venice benchmarks (tomshardware.com)
143.
Waymo says it built a better benchmark for comparing robotaxis to humans (techcrunch.com)
144.
Claude Fable 5 brings Mythos to the masses — Anthropic's new frontier model is 'state-of-the-art on nearly all tested benchmarks' (tomshardware.com)
145.
Benchmarks in Leipzig (news.ycombinator.com)
146.
Unreleased RTX 3050 Ti engineering sample appears in photos and benchmarks — the RTX 3060 alternative that never happened (tomshardware.com)
147.
Chrome for Mac breaks benchmark records on the latest MacBook Pro (9to5mac.com)
148.
Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities (news.ycombinator.com)
149.
These LLMs are the best at resisting Russian propaganda (arstechnica.com)
150.
Benchmark raises its first-ever growth fund as part of $2B capital haul (techcrunch.com)
Today's top topics: polymarket
View all today's topics →