Open-source tool Heretic strips built-in guardrails from AI language models
Heretic is a new tool designed to remove safety restrictions and refusal behaviors from language models, allowing them to follow user instructions without the typical content filters. It targets the alignment layers that model developers add to prevent certain outputs, effectively 'uncensoring' these systems for users who run them.