The fragile foundations of CoT monitoring
(news.ycombinator.com)
1.
2.
Mechanistic interpretability researchers applying causality theory to LLMs
(news.ycombinator.com)
3.
The Download: a new Christian phone network, and debugging LLMs
(technologyreview.com)
4.
This startup’s new mechanistic interpretability tool lets you debug LLMs
(technologyreview.com)
5.
The hidden costs of ‘helpful’ AI
(feeds.nature.com)
6.
Intuitions for Tranformer Circuits
(news.ycombinator.com)
7.
Can you reverse engineer our neural network?
(news.ycombinator.com)
8.
Guide Labs debuts a new kind of interpretable LLM
(techcrunch.com)
9.