Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
91.
Discovering hard disk physical geometry through microbenchmarking (2019) (news.ycombinator.com)
92.
Show HN: A new benchmark for testing LLMs for deterministic outputs (news.ycombinator.com)
93.
A Decade of AMD Ryzen: 10 Years of CPUs Tested (techspot.com)
94.
A Decade of AMD Ryzen: 10 Years of CPUs Tested (techspot.com)
95.
A good AGENTS.md is a model upgrade. A bad one is worse than no docs at all (news.ycombinator.com)
96.
Show HN: OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview (news.ycombinator.com)
97.
Startup unveils benchtop metal 3D printer that brings industrial tech below $10,000 (techspot.com)
98.
The predictable failure of the QDay Prize (news.ycombinator.com)
99.
SWE-bench Verified no longer measures frontier coding capabilities (news.ycombinator.com)
100.
Why SWE-bench Verified no longer measures frontier coding capabilities (news.ycombinator.com)
101.
Mowing Down Simulated Elephants Could Help Self-Driving Cars Prepare For the Chaos of Real Life Streets (futurism.com)
102.
Lambda Calculus Benchmark for AI (news.ycombinator.com)
103.
Linux 7.1 Removes Drivers for Bus Mouse Support (news.ycombinator.com)
104.
OpenAI's GPT-5.5 is here, and it's no potato: narrowly beats Anthropic's Claude Mythos Preview on Terminal-Bench 2.0 (venturebeat.com)
105.
AMD Ryzen 9 9950X3D2 review: More cache, more cash (tomshardware.com)
106.
Kimi vendor verifier – verify accuracy of inference providers (news.ycombinator.com)
107.
Arc Prize Foundation (YC W26) Is Hiring a Platform Engineer for ARC-AGI-4 (news.ycombinator.com)
108.
Experience vs specs: Our readers have spoken, and benchmarks aren’t everything (androidauthority.com)
109.
Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7 (news.ycombinator.com)
110.
Databricks tested a stronger model against its multi-step agent on hybrid queries. The stronger model still lost by 21%. (venturebeat.com)
111.
Databricks research shows multi-step agents consistently outperform single-turn RAG when answers span databases and documents (venturebeat.com)
112.
N-Day-Bench – Can LLMs find real vulnerabilities in real codebases? (news.ycombinator.com)
113.
Android now stops you sharing your location in photos (news.ycombinator.com)
114.
Why we spent 50+ hours retesting Intel’s Core Ultra 270K Plus and 250K Plus (tomshardware.com)
115.
Exploiting the most prominent AI agent benchmarks (news.ycombinator.com)
116.
How We Broke Top AI Agent Benchmarks: And What Comes Next (news.ycombinator.com)
117.
AI models are terrible at betting on soccer—especially xAI Grok (arstechnica.com)
118.
Nubia defends the ethics of REDMAGIC 11 Pro benchmark manipulation (androidauthority.com)
119.
Geekbench 6.7 adds Intel BOT detection to spoof out 'unrealistic' CPU scores — Benchmark runs with BOT enabled will be marked as invalid (tomshardware.com)
120.
Astropad unveils Workbench for Mac: ‘Remote desktop made for the AI era’ (9to5mac.com)
Today's top topics: home abode garage security google
View all today's topics →