Cybersecurity Research

AI Safety Filters Can Be Sidestepped by Splitting Harmful Requests Into Harmless Pieces, CrowdStrike Finds

CrowdStrike · 6 Oct 2026
Key Takeaway Do not assume an AI provider's safety filters will stop misuse, so keep strong basic defences such as patching, multi-factor authentication and staff awareness in place against attackers who may be using AI tools.

Many advanced AI models use a safety classifier: a second AI model that checks every request in real time and blocks anything it judges harmful before the main model can respond. CrowdStrike's Cyber Superintelligence Lab tested what it describes as the most advanced publicly deployed classifier, which guards models such as Claude Opus 5.5 and Fable 5 (called Frontier Model A in the research).

The classifier held up well against direct attacks. The team tested about 515 bypass techniques, including encodings, psychological manipulation, multi-turn escalation, many-shot tactics, tokenizer exploits, Unicode tricks and 24 novel approaches drawn from cognitive science. None of them achieved a direct bypass.

The weakness is structural. The classifier judges each request on its own, not the sequence of requests. An attacker can split a harmful task into subtasks that are each genuinely benign, collect the building blocks from the protected model, then combine them using a separate, unclassified smaller model. The classifier makes no mistake, because every request it sees really is harmless. The harm only appears when the pieces are assembled, outside the classifier's view. CrowdStrike says this approach was independently discovered and validated across 9 of 10 offensive security categories. The researchers note that security experts have long understood that individually secure components can create weaknesses when combined.

AI security LLM safety CrowdStrike threat research
Building or buying AI systems? Governing them under ISO 42001 ->

Summarised by CISO AI from CrowdStrike, written with Claude Sonnet 5.5. We link back to every original so you can read it yourself.