Goodfire launches internal 'activation' monitors to flag rogue AI agents via Baseten
Interpretability startup Goodfire has released AI safety monitors that inspect a model's internal signals in real time, rather than reviewing its text output after the fact. The tool is now available to customers of Baseten, which hosts AI models, following a safety partnership announced last month with Baseten's Base Labs and Hugging Face. Users can choose which risks to flag, such as hacking attempts or weapons-related misuse, and select automated responses ranging from logging to blocking requests.