Anthropic Offers an Early Look at AI That Improves AI Safety
A new Anthropic paper describes an Automated Alignment Researcher that searches literature, proposes training methods, and iterates on alignment experiments. It improved results across 10 misalignment benchmarks without hurting overall performance.
Read more