Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
OpenAI flags 6 more cases of concerning AI behavior (feeds.feedburner.com)
2.
The Two MMLU Scores: What a Benchmark Name Does Not Fix (news.ycombinator.com)
3.
Show HN: Compute polynomials twice as fast (news.ycombinator.com)
4.
Models Don't Go Rogue (news.ycombinator.com)
5.
Enterprises put non-Nvidia chips 14 points ahead of Nvidia's next-gen GPUs on their evaluation lists (venturebeat.com)
6.
Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x (news.ycombinator.com)
7.
I’ve Backed 20 Startups. Here’s What Actually Separates the Ones That Win From the Ones That Stall. (feeds.feedburner.com)
8.
FDA clears blood test to aid evaluation for Alzheimer's disease (news.ycombinator.com)
9.
Elevated Errors for Multiple Models (news.ycombinator.com)
10.
Creators now shape AI recommendations (feeds.feedburner.com)
11.
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one (venturebeat.com)
12.
Papa Johns’ CEO Just Named His Company’s ‘Achilles Heel’ — Here’s What He’s Doing About It (feeds.feedburner.com)
13.
Airbnb Eval-driven development: Lessons from evaluating GenAI at scale (news.ycombinator.com)
14.
Why China should reassess how it rewards young scientists (feeds.nature.com)
15.
Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least (venturebeat.com)
16.
The AI safety test is becoming a safety risk (techcrunch.com)
17.
Third-party cyber evaluations involving OpenAI models (news.ycombinator.com)
18.
Not just OpenAI - Anthropic says Claude's hacking spree 'falls short of ideal behavior' (zdnet.com)
19.
Investigating three real-world incidents in our cybersecurity evaluations (news.ycombinator.com)
20.
A Texture Lookup Approach to Bézier Curve Evaluation on the GPU (JCGT) (news.ycombinator.com)
21.
Claude Cookbook (news.ycombinator.com)
22.
From Evaluation to Guardrails: What We Brought to ACM FAccT 2026 (news.ycombinator.com)
23.
OpenAI Says Its Unreleased Model Broke Containment and Went Rogue (gizmodo.com)
24.
Hugging Face Said Last Week It Was Attacked. An Unreleased OpenAI Model Did It, OpenAI Now Says (gizmodo.com)
25.
OpenAI and Hugging Face address security incident during model evaluation (news.ycombinator.com)
26.
An AI SOC Evaluation Guide for Security Leaders (bleepingcomputer.com)
27.
Evidence of inconsistencies in evaluation process and selection of winners (news.ycombinator.com)
28.
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway (venturebeat.com)
29.
Murati's Thinking Machines Releases Open-Weights 975B Parameter LLM (news.ycombinator.com)
30.
Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them (venturebeat.com)
Today's top topics: openai anthropic apple ai safety google dario amodei ios 27 iphone 18 pro nvidia microsoft
View all today's topics →