An Anthropic researcher just gave us a peek at self-improving AI
By Pradhyuman,
Anthropic published a research paper by Chen Yueh-Han, a researcher in the company's fellows program. The paper, titled Automated Researchers Can Reliably Mitigate Alignment Failures, details how automated systems can improve an artificial intelligence model's performance on alignment benchmarks. During tests on 10 benchmarks for misaligned behaviors, the automated system improved performance on every benchmark. The system did not degrade the overall performance of the model.
The automated system searches literature and proposes a training method. It trains the model for 30 minutes using that method. The paper states that the automated system beats methods proposed by humans within six hours. Anthropic reported that the automated system costs four dollars per hour in API inference. Human researchers cost 150 dollars per hour.
The paper notes that the system depends on benchmarks that accurately reflect alignment goals.
Maintained by Pradhyuman.
Filed under: Anthropic, Models & Research