Anthropic researchers have tested Claude as an automated alignment researcher, allowing it to develop and evaluate methods for addressing 10 types of AI safety failures. Claude closed up to 96 percent of the measured safety gap and achieved an average of 85 percent on deception tests. The researchers also used the weaker Claude Sonnet 5 to improve alignment in an early Opus 4.8 model. The results suggest AI could eventually handle more parts of AI research, although the experiment falls short of true self-improvement.
More in Gadgets
Wozoyo Pure Air PA1 Review: Compact Budget Air Purifier With a Few Missing Features
September 5, 2026
iPhone Ultra Roundup: Launch Date, Expected Price in India, Features, Specifications and More
September 5, 2026
LG Xboom Bounce Review: Looks Like a Sci-Fi Prop, Sounds Surprisingly Good
September 5, 2026
How to Buy Refurbished Phones Safely in India?
September 5, 2026
iPhone 18 Pro Roundup: Expected Price, Launch Date, Specifications, Features and More
September 5, 2026
iPhone 18 Pro Max Roundup: Launch Date, Expected Price in India, Features, Specifications and More
September 5, 2026
Latest News
All you need to know about the Kia Sorento
September 6, 2026
Pandey, Harris and Yastika give Knight Riders opening-night win
September 6, 2026
3 reasons to buy the Lexus ES 500e and 2 to skip it
September 6, 2026
இந்தியாவுக்கு எதிராக பாகிஸ்தான் மகளிர் அணி செய்த படுமோசமான சாதனை.. வரலாற்றிலேயே குறைந்த ஸ்கோர்
September 6, 2026
"பெரிய ஸ்கோர் எடுக்கலாம்னு நினைத்தோம் ஆனால்..".. இந்தியாவிடம் வீழ்ந்த பின் பாகிஸ்தான் கேப்டன் பேச்சு
September 6, 2026
Creative Arts Emmys, Night One: Winners List
September 6, 2026