OpenAI flags 6 more cases of concerning AI behavior
(feeds.feedburner.com)
1.
2.
The Two MMLU Scores: What a Benchmark Name Does Not Fix
(news.ycombinator.com)
3.
Show HN: Compute polynomials twice as fast
(news.ycombinator.com)
4.
Models Don't Go Rogue
(news.ycombinator.com)
5.
Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
(news.ycombinator.com)
6.
7.
Elevated Errors for Multiple Models
(news.ycombinator.com)
8.
Creators now shape AI recommendations
(feeds.feedburner.com)
9.
10.
Airbnb Eval-driven development: Lessons from evaluating GenAI at scale
(news.ycombinator.com)
11.
Why China should reassess how it rewards young scientists
(feeds.nature.com)
12.
13.
14.
Investigating three real-world incidents in our cybersecurity evaluations
(news.ycombinator.com)
15.
Claude Cookbook
(news.ycombinator.com)
16.
17.
18.
OpenAI and Hugging Face address security incident during model evaluation
(news.ycombinator.com)
19.
Evidence of inconsistencies in evaluation process and selection of winners
(news.ycombinator.com)
20.
21.
Hazel (YC W24) Is Hiring for Our Largest Government Contract
(news.ycombinator.com)
22.
Why eval startups fail (2025)
(news.ycombinator.com)
23.
Windows NT for GameCube/Wii
(news.ycombinator.com)
24.
25.
A Man Who Reads Books for a Living (One Every Two Days)
(news.ycombinator.com)
26.
I’ve Hired Hundreds of People — Here’s the Trait I Look For Before Anything Else
(feeds.feedburner.com)
27.
Even (very) noisy LLM evaluators are useful for improving AI agents
(news.ycombinator.com)
28.
The worst job interview I ever had
(news.ycombinator.com)
29.
30.
Charity – Categorical programming language (1998)
(news.ycombinator.com)