Tech News
← Home  ·  All topics

Interpretability

2 GoKawiil briefs on this topic

Goodfire launches internal 'activation' monitors to flag rogue AI agents via Baseten

Interpretability startup Goodfire has released AI safety monitors that inspect a model's internal signals in real time, rather than reviewing its text output after the fact. The tool is now available to customers of Baseten, which hosts AI models, following a safety partnership announced last month with Baseten's Base Labs and Hugging Face. Users can choose which risks to flag, such as hacking attempts or weapons-related misuse, and select automated responses ranging from logging to blocking requests.

Anthropic Publishes Early Framework for Reverse-Engineering Transformer Circuits

Anthropic researchers introduce a mathematical approach to mechanistic interpretability, aiming to reverse-engineer the internal computations of transformer language models. Their initial study focuses on small transformers with two layers or fewer that use only attention blocks, deliberately simpler than models like GPT-3, in order to identify basic patterns before tackling larger systems.