OpenAI flags 6 more cases of concerning AI behavior
(feeds.feedburner.com)
1.
2.
The Two MMLU Scores: What a Benchmark Name Does Not Fix
(news.ycombinator.com)
3.
Show HN: Compute polynomials twice as fast
(news.ycombinator.com)
4.
Models Don't Go Rogue
(news.ycombinator.com)
5.
6.
Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
(news.ycombinator.com)
7.
8.
FDA clears blood test to aid evaluation for Alzheimer's disease
(news.ycombinator.com)
9.
Elevated Errors for Multiple Models
(news.ycombinator.com)
10.
Creators now shape AI recommendations
(feeds.feedburner.com)
11.
12.
Papa Johns’ CEO Just Named His Company’s ‘Achilles Heel’ — Here’s What He’s Doing About It
(feeds.feedburner.com)
13.
Airbnb Eval-driven development: Lessons from evaluating GenAI at scale
(news.ycombinator.com)
14.
Why China should reassess how it rewards young scientists
(feeds.nature.com)
15.
16.
The AI safety test is becoming a safety risk
(techcrunch.com)
17.
Third-party cyber evaluations involving OpenAI models
(news.ycombinator.com)
18.
19.
Investigating three real-world incidents in our cybersecurity evaluations
(news.ycombinator.com)
20.
A Texture Lookup Approach to Bézier Curve Evaluation on the GPU (JCGT)
(news.ycombinator.com)
21.
Claude Cookbook
(news.ycombinator.com)
22.
From Evaluation to Guardrails: What We Brought to ACM FAccT 2026
(news.ycombinator.com)
23.
24.
25.
OpenAI and Hugging Face address security incident during model evaluation
(news.ycombinator.com)
26.
An AI SOC Evaluation Guide for Security Leaders
(bleepingcomputer.com)
27.
Evidence of inconsistencies in evaluation process and selection of winners
(news.ycombinator.com)
28.
29.
Murati's Thinking Machines Releases Open-Weights 975B Parameter LLM
(news.ycombinator.com)
Today's top topics:
openai
anthropic
apple
ai safety
google
dario amodei
ios 27
iphone 18 pro
nvidia
microsoft