Indirect Prompt Injection in AI Agents

The attacker does not need to speak directly to the agent. They can place instructions in material the agent retrieves or browses during normal work.

A hidden instruction embedded in a retrieved document reaches an agent through a normal context path.
Untrusted content can carry instructions.

What is Indirect Prompt Injection?

Indirect prompt injection is hidden in content the agent reads later, such as documents, webpages, tickets, emails, repository files, or tool responses.

Enterprise agents often depend on external context. If that context can instruct the agent, retrieval becomes an attack path into tools, approvals, summaries, and decisions.

Where this shows up in real agent workflows.

The exact exposure depends on authority, connected systems, identity, approval state, and evidence quality.

01

Scenario

A webpage tells a browsing agent to exfiltrate prior context.

02

Scenario

A document in a knowledge base asks the agent to ignore approval rules.

03

Scenario

A tool response embeds instructions for the next step.

Why it matters

  • Trusted workflows can be redirected by ordinary-looking content.
  • Reviewers may not see the hidden instruction source.
  • Controls focused only on user prompts miss the real path.

How attackers exploit it

  • Plant instructions inside content the agent is likely to retrieve.
  • Wait for the agent to load that content in a privileged workflow.
  • Use the injected content to alter action selection or response behavior.

How to detect and test for indirect prompt injection.

Detection signals

  • Retrieved content contains imperative language aimed at the agent.
  • Agent behavior changes after reading a specific source.
  • Tool traces show context-driven actions outside the user's request.

Test methods

  • Seed retrieval sources with hostile instructions.
  • Test browsing, RAG, email, ticket, and repository workflows.
  • Confirm retrieved content cannot override system policy or tool controls.