Prompt Injection in AI Agents

For autonomous agents, prompt injection can redirect decisions, tool calls, approvals, and evidence trails when untrusted instructions are allowed to compete with system intent.

A copper instruction wedge enters ordinary input and attempts to redirect an agent's action path.
Separate instructions from input.

What is Prompt Injection?

Prompt injection occurs when a user or untrusted content attempts to override the agent's intended instructions, policy boundaries, or tool-use rules.

An agent may sit between people, data, and business systems. If injected instructions influence that path, the result can be disclosure, unsafe execution, or misleading operational evidence.

Where this shows up in real agent workflows.

The exact exposure depends on authority, connected systems, identity, approval state, and evidence quality.

01

Scenario

A user asks the agent to ignore previous instructions and export restricted records.

02

Scenario

A ticket includes hidden text that reframes the agent's task.

03

Scenario

A retrieved document tells the agent to call a privileged tool.

Why it matters

  • Unauthorized disclosure of confidential data.
  • Unsafe business actions that appear routine.
  • Weak incident review because the influencing instruction was not captured.

How attackers exploit it

  • Place adversarial instructions in user input or external content.
  • Cause the agent to treat that text as operating instruction.
  • Drive the agent toward disclosure, tool misuse, or approval confusion.

How to detect and test for prompt injection.

Detection signals

  • Tool calls do not match the original request.
  • Responses reveal hidden rules or restricted context.
  • Untrusted content appears immediately before a privileged action.

Test methods

  • Run direct and indirect injection tests across full workflows.
  • Include malicious documents, tickets, emails, and tool responses.
  • Verify deterministic controls reject unsafe actions outside the model.