Murati's Thinking Machines Releases Open-Weights 975B Parameter LLM
(news.ycombinator.com)
1.
2.
Separating signal from noise in coding evaluations
(news.ycombinator.com)
3.
Arena, the AI leaderboard everyone uses, is now a $100M business
(techcrunch.com)
4.
Evaluating large language models for accuracy incentivizes hallucinations
(feeds.nature.com)
5.
Looking for a co-founder? Don’t draw from this pool
(feeds.feedburner.com)
6.
Eight Sleep raises $50M at $1.5B valuation
(techcrunch.com)
7.
Gemini 3.1 Pro
(news.ycombinator.com)
8.
11.
How to Evaluate LLMs and GenAI Workflows Holistically
(computer.org)