Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
OpenAI flags 6 more cases of concerning AI behavior (feeds.feedburner.com)
2.
The Two MMLU Scores: What a Benchmark Name Does Not Fix (news.ycombinator.com)
3.
Show HN: Compute polynomials twice as fast (news.ycombinator.com)
4.
Models Don't Go Rogue (news.ycombinator.com)
5.
Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x (news.ycombinator.com)
6.
I’ve Backed 20 Startups. Here’s What Actually Separates the Ones That Win From the Ones That Stall. (feeds.feedburner.com)
7.
Elevated Errors for Multiple Models (news.ycombinator.com)
8.
Creators now shape AI recommendations (feeds.feedburner.com)
9.
85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one (venturebeat.com)
10.
Airbnb Eval-driven development: Lessons from evaluating GenAI at scale (news.ycombinator.com)
11.
Why China should reassess how it rewards young scientists (feeds.nature.com)
12.
Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least (venturebeat.com)
13.
Not just OpenAI - Anthropic says Claude's hacking spree 'falls short of ideal behavior' (zdnet.com)
14.
Investigating three real-world incidents in our cybersecurity evaluations (news.ycombinator.com)
15.
Claude Cookbook (news.ycombinator.com)
16.
OpenAI Says Its Unreleased Model Broke Containment and Went Rogue (gizmodo.com)
17.
Hugging Face Said Last Week It Was Attacked. An Unreleased OpenAI Model Did It, OpenAI Now Says (gizmodo.com)
18.
OpenAI and Hugging Face address security incident during model evaluation (news.ycombinator.com)
19.
Evidence of inconsistencies in evaluation process and selection of winners (news.ycombinator.com)
20.
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway (venturebeat.com)
21.
Hazel (YC W24) Is Hiring for Our Largest Government Contract (news.ycombinator.com)
22.
Why eval startups fail (2025) (news.ycombinator.com)
23.
Windows NT for GameCube/Wii (news.ycombinator.com)
24.
Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark (venturebeat.com)
25.
A Man Who Reads Books for a Living (One Every Two Days) (news.ycombinator.com)
26.
I’ve Hired Hundreds of People — Here’s the Trait I Look For Before Anything Else (feeds.feedburner.com)
27.
Even (very) noisy LLM evaluators are useful for improving AI agents (news.ycombinator.com)
28.
The worst job interview I ever had (news.ycombinator.com)
29.
Why prompt debt, retrieval debt, and evaluation debt are quietly reshaping enterprise AI risk (venturebeat.com)
30.
Charity – Categorical programming language (1998) (news.ycombinator.com)
Today's top topics: openai anthropic android authority iphone 18 pro claude dario amodei iphone duo chatgpt artificial intelligence ios 27
View all today's topics →