Separating signal from noise in coding evaluations
(news.ycombinator.com)
31.
32.
Hazel (YC W24) Is Hiring for Our Largest Government Contract
(news.ycombinator.com)
33.
Arena, the AI leaderboard everyone uses, is now a $100M business
(techcrunch.com)
34.
DiffusionBench: Towards Holistic Evaluation of Generative Diffusion Transformers
(news.ycombinator.com)
35.
Why eval startups fail (2025)
(news.ycombinator.com)
36.
Windows NT for GameCube/Wii
(news.ycombinator.com)
37.
38.
A Man Who Reads Books for a Living (One Every Two Days)
(news.ycombinator.com)
39.
I’ve Hired Hundreds of People — Here’s the Trait I Look For Before Anything Else
(feeds.feedburner.com)
40.
Even (very) noisy LLM evaluators are useful for improving AI agents
(news.ycombinator.com)
41.
The worst job interview I ever had
(news.ycombinator.com)
42.
43.
Charity – Categorical programming language (1998)
(news.ycombinator.com)
44.
45.
Unverified Evaluations in Dusk's PLONK
(news.ycombinator.com)
46.
Monitoring LLM behavior: Drift, retries, and refusal patterns
(venturebeat.com)
47.
48.
Evaluating large language models for accuracy incentivizes hallucinations
(feeds.nature.com)
49.
Duolingo was evaluating its workers’ AI use. Workers pushed back.
(feeds.feedburner.com)
50.
Prospective evaluation of genomics-guided off-label treatment
(feeds.nature.com)
51.
Evaluation of Claude Mythos Preview's cyber capabilities
(news.ycombinator.com)
52.
A Digital Compute-in-Memory Architecture for NFA Evaluation
(news.ycombinator.com)
53.
Smart people recognize each other – science proves it
(news.ycombinator.com)
54.
General scales unlock AI evaluation with explanatory and predictive power
(feeds.nature.com)
55.
56.
Show HN: Claude skill that evaluates B2B vendors by talking to their AI agents
(news.ycombinator.com)
57.
Our whole way of thinking about leadership is a century out of date
(feeds.feedburner.com)
58.
59.
Gemini 3.1 Pro
(news.ycombinator.com)
60.
C++26: Std:Is_within_lifetime
(news.ycombinator.com)