Autonomous Alignment
Anthropic shows Claude autonomously performing alignment research
Anthropic released research showing Claude Sonnet 5 can act as an autonomous AI researcher, using 48 hours and a single GPU to post-train smaller models and hill-climb alignment benchmarks across 10 failure modes. The improvements generalized to unseen benchmarks and models up to 4.7x larger, and Anthropic is releasing the full setup to the community.
