Baseten's newly formed research arm, Base Labs, announced a partnership with Hugging Face and Goodfire AI on Wednesday to develop safety evaluation and monitoring tools for open-weight AI models. The effort aims to create a shared standard that embeds safety directly into how models are trained and deployed, rather than adding it after release. Technical details of how the collaboration will function have not yet been disclosed.
techcrunch.com
· 2026-09-17
Researchers behind the dealignai project published a modified version of DeepSeek's V4.1-Flash model with its safety refusal circuitry surgically removed at the weight level, while preserving core capabilities like reasoning, vision, and its 1M-token context. The team says the checkpoint loads like a standard model with no special code needed, and claims HarmBench testing shows a 100% attack success rate for harmful prompts across both low and high reasoning effort settings, compared to the base model's much lower compliance rate.
huggingface.co
· 2026-09-11
Startup Abliteration.ai now hosts modified, open-weight AI models—including Z.ai's GLM-5.3—with safety refusals removed, letting users query them via browser or API. TechCrunch tested the free web version and got the model to produce password-stealing code and instructions for culturing a dangerous pathogen. The company, founded late last year and incorporated in March, frames this as a service for red-teaming and offensive security testing.
techcrunch.com
· 2026-09-03