AI Models Are Starting to Train Other AI Systems, Anthropic Study Suggests
Artificial intelligence research may be entering a new phase as AI systems become increasingly capable of helping researchers improve other models. A recent study from Anthropic has provided early evidence that automated AI researchers can make meaningful improvements to the safety and alignment of AI systems.
The development has attracted attention because the ability of AI systems to contribute to the training and improvement of other AI models is often discussed as an important step toward artificial general intelligence (AGI).
AGI generally refers to a hypothetical form of artificial intelligence capable of performing a broad range of tasks at a level that matches or exceeds humans. While Anthropic's research does not establish that AGI has been achieved, the findings suggest that AI-assisted research could become increasingly important in the development of future models.
Anthropic Tests AI Researchers That Train AI Models
The study, titled “Automated Researchers Can Reliably Mitigate Alignment Failures,” was published on August 28. It was led by Chen Yueh-Han, a researcher associated with Anthropic's fellows programme.
Researchers created what they called Automated Alignment Researchers (AARs). These AI-based systems were designed to perform several tasks that would normally require human researchers.
The automated researchers were given access to scientific literature, asked to develop possible training approaches and then used those approaches to improve a target AI model.
Rather than attempting to solve every safety issue at once, the experiment focused on addressing individual alignment failures. The resulting systems were then tested across 10 different benchmarks designed to measure specific types of undesirable or misaligned behaviour.
According to the study, the automated systems improved performance across all of the tested benchmarks without causing an overall decline in the model's performance.
AI Systems Were Able to Search, Experiment and Improve
One of the more notable aspects of the experiment was the amount of research work delegated to the AI systems.
An AAR could examine existing research, suggest a potential training technique and then apply that method to the target model. The training process ran for roughly 30 minutes on an Nvidia H200 GPU during individual iterations.
The researchers retained approaches that produced useful results while removing methods that failed to deliver improvements. This created a process through which successful techniques could progressively be used for further experimentation.
The approach could potentially allow AI-assisted research to operate faster and at a much larger scale than traditional human-only experimentation.
Automated Researchers Also Showed a Cost Advantage
Anthropic also compared the results produced by its automated researchers with approaches suggested by experienced human AI researchers.
The study reported that the strongest automated approach was able to outperform the methods proposed by human researchers, on average, within approximately six hours.
Cost was another significant difference.
Anthropic estimated that its automated researchers cost around $4 per hour in API inference, compared with approximately $150 per hour for the human researchers involved in the comparison.
If similar results can be reproduced in broader settings, AI-assisted research could potentially make certain parts of model development considerably faster and less expensive.
Does This Mean AI Is Moving Toward AGI?
The findings have naturally raised questions about recursive self-improvement—the idea that AI systems could contribute to improving the systems that eventually help create even more capable AI.
However, the Anthropic study should not be interpreted as proof that AI has reached AGI.
The experiment was narrowly focused on alignment research and relied on predefined benchmarks. The automated systems were successful within the environment and objectives established by the researchers, rather than independently deciding what makes an AI system safe or useful in the real world.
This distinction is important because achieving strong results on a benchmark does not necessarily mean an AI system has solved the broader challenges associated with AI alignment.
Human Researchers Are Still Essential
Anthropic's research also highlights several limitations.
Automated alignment researchers depend heavily on the quality of the benchmarks used to evaluate them. If those benchmarks fail to represent real-world alignment objectives, an AI system could appear successful while still missing important safety problems.
Developing, testing and maintaining meaningful benchmarks therefore remains a major responsibility for human researchers.
The AI systems also rely on the existing body of scientific literature. Expanding that research base and discovering entirely new areas of knowledge will continue to require human involvement.
In other words, the study points toward AI working alongside researchers, rather than immediately replacing them.
What Could Come Next?
The significance of the research may ultimately depend on whether similar techniques can be expanded beyond controlled experiments.
If AI systems become capable of conducting increasingly sophisticated experiments, evaluating results and discovering better training methods with limited human intervention, the process of AI development could change substantially.
For now, Anthropic's findings represent an early indication that automated AI research can be useful for specific alignment problems. They also demonstrate why the development of increasingly capable AI systems is being closely watched by researchers working on both performance and safety.
The road to AGI remains uncertain, but studies like this suggest that AI may increasingly become a participant in the research process that creates the next generation of AI itself.

0 Comments