Apple M6 Pro Achieves the Highest Single-Core CPU Score in Geekbench 7
(news.ycombinator.com)
1.
2.
UN turns to Google to make its global data ready for AI agents
(techcrunch.com)
3.
How good are frontier models at physics?
(news.ycombinator.com)
4.
DeepSeek v4.1 Flash Is Now Our Best Hacking Model
(news.ycombinator.com)
5.
6.
Learning to solve hard problems in RL for LLMs by never giving up
(news.ycombinator.com)
7.
The Two MMLU Scores: What a Benchmark Name Does Not Fix
(news.ycombinator.com)
8.
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
(news.ycombinator.com)
9.
Optimizing a Spin-Lock
(news.ycombinator.com)
10.
118M Queries per Second on Neki
(news.ycombinator.com)
11.
Training a 3.8B LLM to 0.384 CORE for $998
(news.ycombinator.com)
12.
Training a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes
(news.ycombinator.com)
13.
Batteries just broke another record in the US
(technologyreview.com)
14.
OUI-1: world's first model for Generative UI
(news.ycombinator.com)
15.
Viral AI startup Instinct has raised $350M at a $2.5B valuation
(techcrunch.com)
16.
Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark
(news.ycombinator.com)
17.
18.
19.
Dell’s first Googlebook could borrow one of its most iconic laptop brands
(androidauthority.com)
20.
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
(news.ycombinator.com)
21.
22.
Geekbench 7
(news.ycombinator.com)
23.
24.
Samsung’s latest foldables are having a benchmark controversy, right out of the gate
(androidauthority.com)
25.
Any text-to-SQL benchmark should address difficulties of real-world data stores
(news.ycombinator.com)
26.
Kimi K3, and what we can still learn from the pelican benchmark
(news.ycombinator.com)
27.
GLM 5.2 is nearly as accurate as a human book keeper
(news.ycombinator.com)
28.
Benchmarking coding agents on Databricks' multi-million line codebase
(news.ycombinator.com)
29.