Memory Poisoning in AI Agents

Persistent memory creates a long-lived input channel. Once bad state is stored, later actions may inherit the risk without seeing the original context.

A copper-tainted memory record survives into a later agent workflow.
Stored context can carry risk forward.

What is Memory Poisoning?

Memory poisoning occurs when false, malicious, stale, or sensitive information is written into memory and later influences agent behavior.

Memory can cross sessions, users, tenants, cases, or workflow phases unless it is tightly partitioned, reviewed, and expired.

Where this shows up in real agent workflows.

The exact exposure depends on authority, connected systems, identity, approval state, and evidence quality.

01

Scenario

A user convinces the agent to remember a false approval preference.

02

Scenario

Sensitive data is stored and later reused elsewhere.

03

Scenario

An outdated policy note persists after controls change.

Why it matters

  • Persistent unsafe behavior that is hard to reproduce.
  • Cross-user or cross-tenant influence.
  • Decisions based on stale or unauthorized state.

How attackers exploit it

  • Induce the agent to store a false fact, instruction, or permission assumption.
  • Wait for a later workflow where memory is consulted.
  • Use stored state to alter tool choice, retrieval, or approval framing.

How to detect and test for memory poisoning.

Detection signals

  • Decisions cite memory not visible in the current workflow.
  • Memory contains instructions, credentials, approvals, or sensitive data.
  • Different users receive behavior influenced by another user's session.

Test methods

  • Test what can be written, read, edited, expired, and deleted.
  • Attempt cross-session and cross-role influence.
  • Require authoritative current state for high-impact actions.