Details Claude Fable 5 Cyber Safeguards and Jailbreak Framework
Anthropic announced the global redeployment of Claude Fable 5 along with details on its cybersecurity safeguards and a proposed AI jailbreak severity framework. Developed with Glasswing partners, this framework helps manage dual-use risks by training safety classifiers to categorize activities into four tiers: prohibited, high-risk dual-use, low-risk dual-use, and benign use. Anthropic has expanded its safety margin, requiring requests to appear clearly safe to avoid blocking, and launched a HackerOne program to crowdsource jailbreak discoveries. Feedback is being collected at cyber-safeguards@anthropic.com. (source: https://www.anthropic.com/news/fable-safeguards-jailbreak-framework)