How can AI agents be attacked?

Updated October 2026 · How we answer

Short answerCommon attacks include prompt injection hidden in web pages or emails, misuse of tools, and leaked credentials. Limiting permissions and treating outside text as data reduces the risk.

Common attack paths

Prompt injection happens when an agent reads text from a web page, email or document that contains hidden instructions. The agent may follow those instructions as if you had written them. Agents with tools are riskier because they can send messages or change files.

Other risks include tools with too much access, leaked API keys and untrusted data sources. Someone who controls one input can sometimes steer the whole workflow.

  • Hidden instructions in web pages or documents
  • Tool permissions that are too broad
  • Leaked keys in logs or code
  • Untrusted data sources

Defenses that help

Give each agent the smallest set of tools and permissions it needs. Require human approval for sending money, deleting data or changing production systems. Treat anything an agent reads from outside as data, not as instructions.

Log tool calls and review them regularly. Test your agent with sample malicious inputs before letting it run unattended. Rotate keys if you suspect they were exposed. Rate limits on outbound actions can also slow down an attack that is already underway.

  • Least-privilege tool access
  • Human approval for risky actions
  • Clear boundaries between instructions and data
  • Log and review tool calls

Common mistakes

  • Trusting text from websites or documents as if it were a command from you.
  • Giving an agent write access to production systems before testing how it behaves.
  • Assuming a model's built-in safety training stops every injected instruction on its own.
From our shopsCaseMorph: Type an idea, see a custom phone case in seconds, then print a one-of-one.