Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
Not just OpenAI - Anthropic says Claude's hacking spree 'falls short of ideal behavior' (zdnet.com)
2.
Investigating three real-world incidents in our cybersecurity evaluations (news.ycombinator.com)
3.
Claude Cookbook (news.ycombinator.com)
4.
OpenAI Says Its Unreleased Model Broke Containment and Went Rogue (gizmodo.com)
5.
Hugging Face Said Last Week It Was Attacked. An Unreleased OpenAI Model Did It, OpenAI Now Says (gizmodo.com)
6.
OpenAI and Hugging Face address security incident during model evaluation (news.ycombinator.com)
7.
Evidence of inconsistencies in evaluation process and selection of winners (news.ycombinator.com)
8.
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway (venturebeat.com)
9.
Murati's Thinking Machines Releases Open-Weights 975B Parameter LLM (news.ycombinator.com)
10.
Separating signal from noise in coding evaluations (news.ycombinator.com)
11.
Hazel (YC W24) Is Hiring for Our Largest Government Contract (news.ycombinator.com)
12.
Arena, the AI leaderboard everyone uses, is now a $100M business (techcrunch.com)
13.
Why eval startups fail (2025) (news.ycombinator.com)
14.
Windows NT for GameCube/Wii (news.ycombinator.com)
15.
Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark (venturebeat.com)
16.
A Man Who Reads Books for a Living (One Every Two Days) (news.ycombinator.com)
17.
I’ve Hired Hundreds of People — Here’s the Trait I Look For Before Anything Else (feeds.feedburner.com)
18.
Even (very) noisy LLM evaluators are useful for improving AI agents (news.ycombinator.com)
19.
The worst job interview I ever had (news.ycombinator.com)
20.
Why prompt debt, retrieval debt, and evaluation debt are quietly reshaping enterprise AI risk (venturebeat.com)
21.
Charity – Categorical programming language (1998) (news.ycombinator.com)
22.
Monitoring LLM behavior: Drift, retries, and refusal patterns (venturebeat.com)
23.
Closure of China’s influential journal ranking leaves academics reeling — what will take its place? (feeds.nature.com)
24.
Evaluating large language models for accuracy incentivizes hallucinations (feeds.nature.com)
25.
Duolingo was evaluating its workers’ AI use. Workers pushed back. (feeds.feedburner.com)
26.
A Digital Compute-in-Memory Architecture for NFA Evaluation (news.ycombinator.com)
27.
Smart people recognize each other – science proves it (news.ycombinator.com)
28.
General scales unlock AI evaluation with explanatory and predictive power (feeds.nature.com)
29.
Duolingo’s CEO Uses a Secret Test to Evaluate Job Candidates — Before They Even Step into the Interview (feeds.feedburner.com)
30.
Show HN: Claude skill that evaluates B2B vendors by talking to their AI agents (news.ycombinator.com)
Today's top topics: apple android app store telegram iphone microsoft universal clipboard openai android authority
View all today's topics →