Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
421.
This iPhone cooling mod turns it into a monstrosity that demolishes benchmarks (androidauthority.com)
422.
When hackathon judging is a public benchmark: my report from Hack the North (news.ycombinator.com)
423.
Snapdragon 8 Elite Gen 5 benchmarks: Just how badly does it beat its rivals? (androidauthority.com)
424.
SWE-Bench Pro (news.ycombinator.com)
425.
Evals in 2025: going beyond simple benchmarks to build models people can use (news.ycombinator.com)
426.
Evals in 2025: benchmarks to build models people can use (news.ycombinator.com)
427.
Tau² benchmark: How a prompt rewrite boosted GPT-5-mini by 22% (news.ycombinator.com)
428.
Tau² Benchmark: How a Prompt Rewrite Boosted GPT-5-Mini by 22% (news.ycombinator.com)
429.
Crowdstrike and Meta just made evaluating AI security tools easier (zdnet.com)
430.
Why do browsers throttle JavaScript timers? (news.ycombinator.com)
431.
Top model scores may be skewed by Git history leaks in SWE-bench (news.ycombinator.com)
432.
DeepCodeBench: Real-World Codebase Understanding by Q&A Benchmarking (news.ycombinator.com)
433.
Pixel 10 benchmarks show how little ground Google’s Tensor G5 has gained (androidauthority.com)
434.
MCP-Universe benchmark shows GPT-5 fails more than half of real-world orchestration tasks (venturebeat.com)
435.
Benchmarks for Golang SQLite Drivers (news.ycombinator.com)
436.
Researchers Solve 35-Year-Old Fusion Mystery With Bench-Top Reactor (gizmodo.com)
437.
Worried about the Pixel 10 Pro XL benchmark controversy? Here’s why you shouldn’t be (androidauthority.com)
438.
Stop benchmarking in the lab: Inclusion Arena shows how LLMs perform in production (venturebeat.com)
439.
Small Objects, Big Gains: Benchmarking Tigris Against AWS S3 and Cloudflare R2 (news.ycombinator.com)
440.
Pixel 10 Pro XL benchmark leaks are making people mad (androidauthority.com)
441.
Show HN: Evaluating LLMs on creative writing via reader usage, not benchmarks (news.ycombinator.com)
442.
GPT-5 Doesn't Dislike You—It Might Just Need a Benchmark for Emotional Intelligence (wired.com)
443.
Launch HN: Design Arena (YC S25) – Head-to-head AI benchmark for aesthetics (news.ycombinator.com)
444.
Qodo CLI agent scores 71.2% on SWE-bench Verified (news.ycombinator.com)
445.
Benchmarking GPT-5 on 400 real-world code reviews (news.ycombinator.com)
446.
Benchmark Framework Desktop Mainboard and 4-node cluster (news.ycombinator.com)
447.
Herbie detects inaccurate expressions and finds more accurate replacements (news.ycombinator.com)
448.
MSI Claw benchmarked using AMD and Intel chips: Ryzen Z2 Extreme trumps Core Ultra 7 258V (techspot.com)
449.
Do LLMs identify fonts? (news.ycombinator.com)
450.
Efficiently Generating a Number in a Range (2018) (news.ycombinator.com)
Today's top topics: android authority artificial intelligence anthropic donald trump openai polymarket
View all today's topics →