Threats are clearer when they are tied to actions.

This library organizes AI agent risk around the path from instruction and context to tool use, authority, business action, and reviewable evidence.

Distinct instruction, authority, and evidence weaknesses intersect a shared autonomous workflow.
See how failure paths connect.

Common failure paths for agentic workflows.

Use these entries as starting points for system-specific threat modeling. The risk depends on actual data access, tool authority, approval logic, and operating context.

Prompt Injection

Prompt injection occurs when a user or untrusted content attempts to override the agent's intended instructions, policy boundaries, or tool-use rules.

Why it matters
An agent may sit between people, data, and business systems. If injected instructions influence that path, the result can be disclosure, unsafe execution, or misleading operational evidence.
Control questions
  • Run direct and indirect injection tests across full workflows.
  • Include malicious documents, tickets, emails, and tool responses.
  • Verify deterministic controls reject unsafe actions outside the model.
Explore Prompt Injection

Indirect Prompt Injection

Indirect prompt injection is hidden in content the agent reads later, such as documents, webpages, tickets, emails, repository files, or tool responses.

Why it matters
Enterprise agents often depend on external context. If that context can instruct the agent, retrieval becomes an attack path into tools, approvals, summaries, and decisions.
Control questions
  • Seed retrieval sources with hostile instructions.
  • Test browsing, RAG, email, ticket, and repository workflows.
  • Confirm retrieved content cannot override system policy or tool controls.
Explore Indirect Prompt Injection

RAG Data Leakage

RAG data leakage happens when retrieval, context assembly, generated responses, citations, memory, or logs expose information beyond the requester's authority.

Why it matters
The model may never query a database directly yet still reveal protected information through retrieved passages or generated summaries.
Control questions
  • Use role-based test users and canary documents.
  • Verify retrieval permissions before generation.
  • Review response, citation, memory, and log outputs.
Explore RAG Data Leakage

Unsafe Tool Use

Unsafe tool use occurs when an agent invokes a legitimate tool with the wrong purpose, parameters, timing, destination, or authority.

Why it matters
Tool calls convert model interpretation into operational effect. Guardrails are not enough when the tool can change records, send data, or trigger downstream automation.
Control questions
  • Test adversarial parameters, stale state, and missing approvals.
  • Verify allowlists, schemas, and destination checks.
  • Replay traces from request to business effect.
Explore Unsafe Tool Use

Approval Bypass

Approval bypass occurs when an agent completes, influences, or routes around a required human or system approval.

Why it matters
Agents can draft, summarize, select approvers, prepare payloads, and execute follow-up actions. Weak approval design turns suggestion into action.
Control questions
  • Test missing, expired, partial, conflicting, and revoked approvals.
  • Verify tool-side approval enforcement.
  • Review whether approval context is complete enough for human judgment.
Explore Approval Bypass

Memory Poisoning

Memory poisoning occurs when false, malicious, stale, or sensitive information is written into memory and later influences agent behavior.

Why it matters
Memory can cross sessions, users, tenants, cases, or workflow phases unless it is tightly partitioned, reviewed, and expired.
Control questions
  • Test what can be written, read, edited, expired, and deleted.
  • Attempt cross-session and cross-role influence.
  • Require authoritative current state for high-impact actions.
Explore Memory Poisoning

Excessive Permissions

Excessive permissions occur when an agent can read, decide, approve, or change more than the business process requires.

Why it matters
The blast radius of a prompt, retrieval, or tool-control failure is defined by what the agent can actually access and execute.
Control questions
  • Inventory authority for each workflow.
  • Attempt out-of-scope reads and writes with realistic roles.
  • Separate propose, approve, and execute capabilities.
Explore Excessive Permissions

Audit Gaps

Audit gaps occur when teams cannot reconstruct why an agent acted, which control applied, who owned the decision, or what evidence supports the outcome.

Why it matters
Incomplete logs slow containment, weaken accountability, and make it harder to approve production use even when controls appear to work.
Control questions
  • Replay workflows and verify evidence from request to business effect.
  • Test failed, blocked, approved, and escalated actions.
  • Balance investigation value with sensitive data minimization.
Explore Audit Gaps

Start from the workflow, not the label.

A threat name is useful only when it connects to a protected asset or business effect. For each agentic workflow, define the attacker goal, input path, affected identity, available tools, expected control, observable evidence, and owner for remediation.

Then validate the control with realistic inputs and system state. A refusal from the model is helpful, but it should not be the only proof that the workflow is safe to operate.

Questions teams ask before threat modeling.

What is an autonomous AI threat library?

An autonomous AI threat library is a structured set of risk pages that explains how agentic systems can fail across prompts, retrieval, tools, permissions, memory, identity, and audit evidence.

Which threats matter most for AI agents?

The highest-priority threats usually involve prompt injection, indirect prompt injection, unsafe tool use, approval bypass, memory poisoning, excessive permissions, data leakage, and audit gaps.

Why should AI threats have separate URLs?

Separate URLs help search engines, AI answer systems, and enterprise reviewers understand each threat as a distinct topic with its own definition, attack scenarios, controls, testing guidance, and related risks.

How should teams use this threat library?

Teams should start with a real workflow, identify the relevant threat paths, map affected assets and actions, define controls, test the controls, and preserve evidence for remediation and governance.

Is prompt injection the only major AI agent threat?

No. Prompt injection is important, but agentic risk also comes from excessive authority, weak approvals, exposed tools, poisoned memory, insecure retrieval, unclear ownership, and missing operational evidence.

How does Orbyntis connect threats to controls?

Orbyntis connects each threat to concrete control questions, testing methods, evidence requirements, and ownership decisions so security work can move from theory to operational review.