Anthropic fellow demonstrates AI system that autonomously fixes model alignment flaws
An Anthropic fellows program researcher, Chen Yueh-Han, published a paper showing an automated system that searches literature, proposes fixes, and trains models to improve performance across 10 alignment benchmarks without hurting overall capability. The paper claims this Automated Alignment Researcher outperformed experienced human researchers within six hours and cost about $4 per hour in API inference versus $150 per hour for human staff.