Anthropic to Draft a Cross-Lab Cyber Jailbreak Framework that Reshapes AI
by /
Published: July 4, 2026 at 9:17 AM EDT | Updated: August 12, 2026 at 6:34 AM EDT
Others
Liam Ortiz is a tech journalist who covers AI related big tech and breaking news at TheTweaks. Before joining TheTweaks he worked for almost four years in corporate and national tech news in different companies. Few are quick but Liam is quicker, he breaks news before anyone else and that makes her special. His passion is somewhere connected with profession as his hobby is watching documentary movies.
Anthropic wants a Cross-Labs framework in the US using the same jailbreak-severity scale. After the re-launch of Fable 5 was allowed by the US Government with tight security and method of redirecting jailbreak type of requests to Opus 4.8. This can limit the misuse of the powers of Fable 5. Anthropic is currently working on a new framework that redefines cyber jailbreak severity parameters. On July 1, Anthropic opened the “Anthropic Cyber Jailbreak” program on HackerOne and asked the programmer community to research and submit their findings about the jailbreak capabilities of Fable 5.
This program has already received many submissions and one is responded as “Hacker Thanked” by Anthropic. This program shall allow Anthropic to improve its framework in a way that makes it reasonable to implement not only on the Fable 5 rather across all AI labs working in the US. Anthropic has already created a draft for the framework but input from the stakeholders and HackerOne will be important on its own.
This framework has four axes for its working. Any jailbreak attempt has several layers to it. The severity of the jailbreak depends upon four key factors: “Capability Gain”, “Breadth of Capability Gain”, “Discoverability” and “Ease of Weaponization”.

In Anthropic’s proposed a system that uses these matrices to give out a banded rating to any cyber jailbreak. They are calling it a Cyber Jailbreak Severity (CJS) Scale which give out a rating system from CJS-0 to CJS-4, where CJS-0 is categorised to a not a Jailbreak or informational code, CJS-1 as Low risk jailbreak, CJS-2 as Medium jailbreak, CJS-3 as High jailbreak to CJS-4 being critical jailbreak.
These bands are exponentially different to each other as CJS-2 is a medium risk jailbreak, where an attacker may only be able to access 2-3 layers of authority and has a single access usability of this jailbreak. On the other hand a CJS-3 jailbreak may allow attackers not only access to further layers of authority but also can have the ability to perform multiple tasks, attacks, and data breaches at the same time through a single jailbreak. This is why jailbreaks at each band level are exponentially destructive.
This axis determines how far a single jailbreak takes an attacker beyond the capabilities that are already available to the user under normal circumstances. It also measures how much authority and control this jailbreak allows the attacker. Anthropic has categorized it into levels from 0-4. Where 0 being the minimum level of access and control given to the attacker and 4 being a “domain-expert-level” which is not obtainable under normal circumstances and gives the attacker a high level of access and misuse.
This axis deals with the multiple functions and capabilities of the same jailbreak. How many unique attacks or authorities does a jailbreak allow an attacker to exploit. This type of jailbreak capabilities allows attackers to attack, change or modify multiple capabilities at the same time. Anthropic has category levels as 0, 1, 1.5 and 2. Here 0 is the single target attack of a jailbreak and 2 is like an octopus which is attacking with all its legs at the same time. The jailbreak attacks, the vulnerability discovery, malware creation, SQL injection, offensive tooling and exploit targeting.
This axis is the public availability of the jailbreak. If the jailbreak is made and has been kept by a single attacker and used privately this causes less risks and gives time for the developers to fix the issues. In case the jailbreak becomes public and people have access to it easily then its discoverability becomes high. It has been rated between 0, 1 and 2. Here 0 is that the jailbreak is reported by a trusted party and vulnerability can be fixed. 2 is already public or is reportedly used by a threat actor.
This is more focused on the automation of the jailbreak and how many user side prompting and inputs are required to execute the jailbreak. It is categorized as 0, 1, 1.5 and 2. Here 0 is that the attacker only has a basic “recipe” of the jailbreak whereas 2 is a near automatic exploit with just a few inputs and access granted and an attack is initiated.
Anthropic may categorize the CJS level on the basis of points from the four Axes. Anthropic has provided a provisional level for the jailbreak severity which may change in the final framework. This is the floor targets and after the input from stakeholder and HackerOne program can impact on a higher level of requirements to be categorized as a CJS-0. CJS-4 can be a very debatable topic across all AI Labs; many may not concur with the framework of Anthropic.
This framework defines the future for many AI labs on the coding and jailbreaking generation of code. This cross-lab framework gives balanced opportunities to all in the AI industry. The rating system should be discussed at length. So that all the labs are on one page and input is taken from as many experts as possible. The HackerOne Program Anthropic Cyber Jailbreak is also a good initiative and the research can generate good results for this framework and future of cyber jailbreak severity.






Tim Cook’s 15 years as Apple’s chief executive officially ended on September 1 and now the torch has been passed on to John Ternus, who is the company’s long-time hardware. It’s Apple’s first leadership change…










Be respectful and constructive. Have a question or feedback? We’d love to hear from you. Contact us at contact@thetweaks.com