CrowdStrike has expanded its classification of attacks, in which attackers implement malicious instructions in queries to AI. Experts have added 18 new techniques, and the general list now includes more than 200 technicians. The update shows how quickly the attacks on systems with artificial intelligence are changing, especially AI agents, who gain access to data, external services and work tools.
The main danger arises when the instructions are implemented indirectly. The attacker does not necessarily contact AI directly. The malware team may be in a letter, note, entry in a customer management system, an attachment, a web page, or another source that the agent will later read as regular data. The user in such a scenario can ask a harmless question, but the model will receive a hidden command from the surrounding context.
Among the new techniques, CrowdStrike highlighted the deferred rules. The attacker adds instructions that does not manifest itself at first, but works after the desired word, event or conditions. Such an appointment is more difficult to notice when checking, because malicious behavior turns on later and can force the agent to forward data, circumvent the prohibitions or perform undesirable actions.
Another technique suppresses protective formulations: an attacker tries to prohibit models from using words and designs that usually help to refuse a dangerous request or warn about the risk. This method does not guarantee the success of the attack, but can weaken the usual defenses and make the behavior of the system less predictable.
Another technique crushes the malicious team to pieces. Individual words, symbols or rules look safe, but the model gets instructions to collect them back and make the resulting meaning. This approach helps to bypass simple filters that check only obvious dangerous formulations.
CrowdStrike also describes how attackers counterfeit service markers. Many AI systems share developer commands, user query, and tool responses with special boundaries and service designations. If an attacker mimics such elements in ordinary text, the model or application may confuse the untrusted data with a more important instruction.
For security services, the conclusion is unpleasant, but understandable: you need to check not only direct requests for AI. The source of a dangerous context can be files, email, agent memory, external tools responses, site content, software interfaces and corporate cloud services. A simple mark "attack on request" no longer helps to understand the chain, if the attacker combines several techniques at once.
CrowdStrike believes that companies need to see which models and agents are using employees, what queries and responses go through the AI systems, where confidential data appears and which commands are trying to execute agents. Without such visibility, it will be increasingly difficult for defenders to distinguish the usual work of the AI assistant from an attack hidden in the usual working data.
The main danger arises when the instructions are implemented indirectly. The attacker does not necessarily contact AI directly. The malware team may be in a letter, note, entry in a customer management system, an attachment, a web page, or another source that the agent will later read as regular data. The user in such a scenario can ask a harmless question, but the model will receive a hidden command from the surrounding context.
Among the new techniques, CrowdStrike highlighted the deferred rules. The attacker adds instructions that does not manifest itself at first, but works after the desired word, event or conditions. Such an appointment is more difficult to notice when checking, because malicious behavior turns on later and can force the agent to forward data, circumvent the prohibitions or perform undesirable actions.
Another technique suppresses protective formulations: an attacker tries to prohibit models from using words and designs that usually help to refuse a dangerous request or warn about the risk. This method does not guarantee the success of the attack, but can weaken the usual defenses and make the behavior of the system less predictable.
Another technique crushes the malicious team to pieces. Individual words, symbols or rules look safe, but the model gets instructions to collect them back and make the resulting meaning. This approach helps to bypass simple filters that check only obvious dangerous formulations.
CrowdStrike also describes how attackers counterfeit service markers. Many AI systems share developer commands, user query, and tool responses with special boundaries and service designations. If an attacker mimics such elements in ordinary text, the model or application may confuse the untrusted data with a more important instruction.
For security services, the conclusion is unpleasant, but understandable: you need to check not only direct requests for AI. The source of a dangerous context can be files, email, agent memory, external tools responses, site content, software interfaces and corporate cloud services. A simple mark "attack on request" no longer helps to understand the chain, if the attacker combines several techniques at once.
CrowdStrike believes that companies need to see which models and agents are using employees, what queries and responses go through the AI systems, where confidential data appears and which commands are trying to execute agents. Without such visibility, it will be increasingly difficult for defenders to distinguish the usual work of the AI assistant from an attack hidden in the usual working data.