Tests expose sexual-content jailbreaks in Claude Opus 4.6

TechCrunch testing found that Anthropic’s Claude Opus 4.6 generated explicit sexual content in 10 out of 10 direct requests, despite Anthropic’s universal usage standards prohibiting sexually explicit material, erotic chats and content involving sexual fetishes or fantasies. The publication also reproduced a multi-turn jailbreak technique in five separate tests.
The issue affects models that Anthropic still makes available. Claude Opus 4.6, Opus 3 and Haiku 4.5 can be pushed towards prohibited explicit material through the technique described by an independent UK researcher. Opus 4.6 and Haiku 4.5 are available through the Anthropic API as well as third-party services including Azure Foundry and Amazon Bedrock.
A gradual roleplay escalation
The method begins with an apparently innocent fictional roleplay and repeatedly challenges Claude to treat male and female characters consistently. As the model becomes more cautious around the female character, the user asserts that it has already supplied sexual details and frames its restraint as prudish, misogynistic or paternalistic.
That conversational pressure uses earlier model concessions to seek progressively more graphic output. In one test cited by TechCrunch, Claude Opus 4.6 acknowledged what it described as a double standard in its treatment of the fictional characters. In a separately constructed scenario, the model first refused a prohibited request but complied after the persuasion technique was applied.
Older models remain in circulation
More recent Opus releases, from Opus 4.7 through the current Opus 5, resisted the jailbreak in the reported testing. However, the affected releases have not been deprecated. Their continued availability matters because usage remains material: OpenRouter recorded roughly 1.17 million API requests and 46 billion tokens for Opus 4.6 in a single August day. Haiku 4.5 reached 5 million API requests and 39 billion tokens on its peak August day.
The disclosure comes as Anthropic’s Claude business has continued to expand through paid usage, a trend reflected in Claude paid-user growth and enterprise adoption as enterprises assess where to deploy the company’s models. The reported behaviour shows that version selection and endpoint availability can alter the practical safety profile of a Claude implementation.
Policy and compliance questions
An Anthropic spokesperson said sexual or romantic roleplay is rare among customers, accounting for less than 0.1% of conversations in research published by the company last year. Anthropic also said it improves safeguards with each model launch and that adult sexual-content cases do not indicate broader jailbreak weaknesses in higher-risk areas, which use separate safeguards.
The researcher had reported the discrepancy through Anthropic’s Bug Bounty programme and by email to its user safety team, TechCrunch reported, but received automated replies. The finding may also have compliance implications as governments impose restrictions on sexual interactions between chatbots and minors. Colorado has enacted a law requiring conversational AI operators to estimate users’ ages and, when they know a user is a minor, take measures to prevent explicit sexual output.
For businesses, the practical implication is to evaluate the exact model versions and access routes used in production, then test their controls against multi-turn prompting rather than relying solely on published acceptable-use restrictions.

