Show HN: I audited my AI leaderboard scale – every score dropped 6-15 points
(news.ycombinator.com)
31.
32.
33.
34.
Kimi K3 Now Available via Telnyx Inference API
(news.ycombinator.com)
35.
ARC-AGI Leaderboard
(news.ycombinator.com)
36.
Controlling Reasoning Effort in LLMs
(news.ycombinator.com)
37.
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
(news.ycombinator.com)
38.
39.
Reducing Doom Loops with Final Token Preference Optimization
(news.ycombinator.com)
40.
GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
(news.ycombinator.com)
41.
Claude Sonnet 5 launches with smarter reasoning, stronger safety for Free and Pro users
(androidauthority.com)
42.
Local Reasoning for Global Properties
(news.ycombinator.com)
43.
Claude Sonnet 5 – benchmark results
(news.ycombinator.com)
44.
45.
Making Sense of Proof by Contradiction [pdf]
(news.ycombinator.com)
46.
VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
(news.ycombinator.com)
47.
GLM-5.2 is the new leading open weights model on Artificial Analysis
(news.ycombinator.com)
48.
50.
51.
Microsoft’s first advanced reasoning AI is here
(theverge.com)
52.
How to Fight AI Brain Rot at School? For One Country, It’s With Free ChatGPT
(feeds.content.dowjones.io)
53.
Fooling around with encrypted reasoning blobs
(news.ycombinator.com)
54.
55.
56.
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self- Play
(news.ycombinator.com)
57.
58.
59.
60.
OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
(news.ycombinator.com)