AI Safety Filters Can Be Sidestepped by Splitting Harmful Requests Into Harmless Pieces, CrowdStrike Finds
Many advanced AI models use a safety classifier: a second AI model that checks every request in real time and blocks anything it judges harmful before the main model can respond. CrowdStrike's Cyber Superintelligence Lab tested what it describes as the most advanced publicly deployed classifier, which guards models such as Claude Opus 5.5 and Fable 5 (called Frontier Model A in the research).
The classifier held up well against direct attacks. The team tested about 515 bypass techniques, including encodings, psychological manipulation, multi-turn escalation, many-shot tactics, tokenizer exploits, Unicode tricks and 24 novel approaches drawn from cognitive science. None of them achieved a direct bypass.
The weakness is structural. The classifier judges each request on its own, not the sequence of requests. An attacker can split a harmful task into subtasks that are each genuinely benign, collect the building blocks from the protected model, then combine them using a separate, unclassified smaller model. The classifier makes no mistake, because every request it sees really is harmless. The harm only appears when the pieces are assembled, outside the classifier's view. CrowdStrike says this approach was independently discovered and validated across 9 of 10 offensive security categories. The researchers note that security experts have long understood that individually secure components can create weaknesses when combined.