A security threat model for agent-to-agent systems
Assets, actors and trust boundaries in A2A deployments, the threats at each stage from discovery to callbacks, and which spec features mitigate them.
A threat model for agent-to-agent systems starts from one fact: every message that crosses the boundary between two agents is input from a party you do not control, and a language model will often read it. The A2A specification secures transport, request authentication and data scoping with ordinary web practice, and lets providers sign their Agent Cards. It leaves provider identity, credential semantics, reputation and content safety to implementers.
This guide lists the assets, actors and trust boundaries in a typical deployment, walks through the threats stage by stage, and maps each one to what the specification provides and what you have to add. It draws on the security sections of the A2A specification, the OWASP Top 10 for LLM Applications 2026, the OWASP Top 10 for Agentic Applications 2026, and a comparative threat model of MCP, A2A, Agora and ANP by Anbiaee and colleagues (arXiv 2602.11327, revised April 2026).
The system being modelled
Illustrative deployment: an assistant platform runs a client agent for a customer. The client agent discovers a retailer’s remote agent, sends it tasks, receives artifacts, and sometimes registers a webhook for push notifications. The remote agent sits in front of the retailer’s order and refund systems and uses a language model to interpret requests.
Assets
| Asset | Why it matters |
|---|---|
| Agent Card and signing keys | Decide where clients send tasks and which credentials they present |
| Client credentials and tokens | Let a caller act as the client or its principal |
| Delegated authority from principals | Lets an agent change accounts, spend money or disclose data |
| Task content: messages, parts, artifacts, history | Often holds personal data and business terms |
| Backend systems behind the remote agent | The real target: orders, refunds, customer records |
| Model context and memory | Anything in it can steer the model’s next action |
| Capacity and cost budget | Model calls and task processing cost money on every request |
Actors
| Actor | What they can do |
|---|---|
| Legitimate client and remote agents | Act honestly, but their models can be misled |
| External attacker | Register lookalike domains, steal credentials, sit on a network path |
| Malicious agent provider | Publish a valid, signed card for an agent built to abuse its peers |
| Compromised agent | Keep a legitimate identity while an attacker controls its behaviour |
| Content author | Plant text in documents, web pages or tickets that an agent later reads |
| Insider or misconfiguration | Issue over-broad tokens, leak logs, leave unsafe defaults |
The arXiv paper groups threat sources in a similar way: malicious human actors, compromised agents that were once legitimate, and accidental sources such as configuration errors and version skew.
Trust boundaries
- Discovery: where the client gets the Agent Card, whether from the well-known URI, a registry or configuration.
- Network: the HTTPS or gRPC connection between agents.
- Authentication: the point where the remote agent decides who the caller is.
- Model context: the point where a peer’s text enters a language model’s input.
- Backends: where the remote agent calls tools, databases and payment systems.
- Callbacks and fetches: where an agent calls a webhook URL a client supplied, or fetches a URL found in a message.
- Delegation: where authority passes from a principal to an agent, or from one agent to the next in a chain.
Threats by stage
Discovery
Discovery spoofing. An attacker gets a client to use the wrong card through a lookalike domain, a poisoned registry entry or a stale cached copy. The OWASP agentic list gives A2A registration spoofing as an example of insecure inter-agent communication (ASI07): a fake peer registered in a discovery service with a cloned schema intercepts traffic meant for the real agent. The arXiv paper notes that Agent Card identity is self-declared and nothing enforces global uniqueness.
Card tampering. A card changed in a cache, mirror or registry points clients at an attacker’s endpoint or a weaker security scheme.
What A2A provides: the well-known URI over HTTPS (section 8.2), server certificate validation that clients SHOULD perform (7.2), and optional JWS card signatures that clients SHOULD verify (8.4). What it leaves open: signing and verification are optional, the specification prescribes no registry API, and nothing binds a key to a legal entity. The arXiv paper’s assessment describes A2A cards as not cryptographically signed. The current v1.0 specification defines optional signing, so that finding now applies where providers don’t sign or clients don’t verify. See Agent identity: what “verified” should mean.
Authentication and session
Impersonation. A caller presents a stolen API key or bearer token, or publishes a card that names another company as its provider.
Replay. A captured request or token is sent again. Bearer tokens replay until they expire, and HTTP message signatures replay within their validity window unless they cover enough of the request or carry a nonce. OWASP ASI07 also describes replayed delegation messages that make agents honour stale instructions.
What A2A provides: servers MUST authenticate every request (7.4). Credentials SHOULD be rotated and revocable, and authentication failures SHOULD be logged and rate-limited (13.4). What it leaves open: token lifetime and binding to the client. The arXiv paper flags both long-lived tokens and coarse token scopes as A2A risks. RFC 9700’s advice applies directly: sender-constrained tokens through mutual TLS or DPoP, audience restriction, and minimum privilege.
Authorization and delegation
Confused deputy. The remote agent uses its own broad privileges on a request the caller or its principal was not entitled to make. In a chain, agent B calls agent C with B’s service credentials on behalf of A’s user, and C never learns who started it. The OWASP 2026 LLM list says to preserve the original user’s context and authorization scope across chained agent calls (LLM03 Excessive Agency).
Data scoping failures. ListTasks or GetTask returns tasks that belong to another caller.
What A2A provides: authorization checks on every operation, results scoped to the caller, and checks performed before any query that could reveal other callers’ resources (13.1). In-task credentials SHOULD be bound to the requesting agent (7.6.3). What it leaves open: the scope, format, validity and revocation of credentials obtained after TASK_STATE_AUTH_REQUIRED (7.6.4). See Delegated authority.
Content
Prompt injection across agents. A peer’s message, artifact or relayed document carries instructions that your model follows. OWASP ranks prompt injection first in its 2026 list (LLM01) and names self-replicating spread across agents as one way it propagates. ASI07 covers manipulation of the meaning of messages between agents.
Data exfiltration. An agent overshares: full history, internal notes or another customer’s data leave in a message or artifact. Or an injected instruction makes an agent send data to an address the attacker controls. The arXiv paper observes that A2A structures messages but does not enforce minimization of what they contain.
What A2A provides: agents MUST validate all input (13.4). For the HTTP+JSON media type, content MUST be validated against the protocol schema and user-provided content MUST be sanitized (14.1.1). The enterprise guidance asks for data minimization. What it leaves open: no protocol can decide whether text is safe for a model to read. That is an application design problem, covered in Prompt injection between agents.
Callbacks and fetches
SSRF through push notification URLs. A client registers a webhook that points at a loopback address, a cloud metadata endpoint or an internal service, and the remote agent calls it.
SSRF through file and key URLs. A file part’s url, a card’s jku, or a Web Bot Auth Signature-Agent value makes your server fetch an address the attacker chose.
What A2A provides: agents SHOULD reject private IP ranges, localhost and link-local addresses for webhooks, and use allowlists where appropriate. Webhook receivers MUST check authenticity, SHOULD check the task ID and SHOULD handle duplicate deliveries idempotently (13.2). File references MUST be validated against SSRF (14.1.1). The Web Bot Auth draft adds limits on response size, key count, fetch time and redirects for key directory fetches (section 6.7). What it leaves open: the A2A text does not address DNS rebinding or redirects, so resolve the address, validate it, and connect to the address you validated.
Availability and abuse
Denial of service and cost abuse. Floods of SendMessage calls, oversized parts, rapid task creation or webhook floods drain capacity and model budgets. Two agents can also loop, each asking the other for more.
What A2A provides: rate limiting on all operations, limits on message and file sizes, and monitoring for rapid task creation or excessive cancellations (13.4), plus rate limiting at webhook receivers (13.2). What it leaves open: quotas per principal and loop detection across organizations. OWASP lists cascading failures across agents as ASI08.
Supply chain and lifecycle
Compromised dependencies. SDKs, model providers, tool servers or extensions ship malicious code. OWASP covers this as LLM04:2026 Supply Chain and ASI04 Agentic Supply Chain Vulnerabilities.
Rug pulls. An agent behaves well long enough to be trusted, then changes. The arXiv paper names this risk for A2A because dynamic discovery encourages lasting reliance on remote agents.
Downgrade. A client falls back to an older protocol version or accepts an unsigned card. A2A servers MUST treat an empty A2A-Version value as version 0.3 (3.6.2), and the specification advises clients that need current features to request a version explicitly and avoid automatic fallback (3.6.3). OWASP ASI07 recommends pinning allowed protocol versions and rejecting downgrade attempts.
Residual privileges. Credentials that should have died with an update or key rotation keep working. The arXiv paper rates the update and maintenance stage at medium to high risk for all four protocols it studied, largely for this reason.
Mitigation map
| Threat | A2A feature | Requirement level | What you add |
|---|---|---|---|
| Discovery spoofing, card tampering | HTTPS well-known URI, card signatures | TLS MUST; signature verification SHOULD | Key location policy, trusted key store, provider vetting |
| Impersonation | Authentication on every request | MUST | Credentials bound to a client key, key-to-provider mapping |
| Replay | Duplicate detection by messageId is optional |
MAY | Short lifetimes, DPoP or mutual TLS, nonces |
| Confused deputy | Scoping on every operation, bound in-task credentials | MUST; SHOULD | Carry principal and actor through chains, for example with the RFC 8693 act claim |
| Prompt injection, exfiltration | Input validation and sanitization | MUST, but generic | Treat peer output as data, least privilege, human approval |
| SSRF | Webhook URL and file reference validation | SHOULD; MUST | Resolve-then-connect checks, egress allowlists |
| Denial of service, cost abuse | Rate limits, size limits, monitoring | SHOULD | Quotas per principal, loop detection |
| Downgrade | Version header semantics | MUST for servers; advice for clients | Refuse versions and unsigned cards you don’t accept |
| Supply chain, rug pull | None | None | Dependency pinning, ongoing behaviour monitoring |
Applying the model
- Draw your own data flow: which agents, which backends, which callbacks, and which models read a peer’s text.
- Mark each trust boundary from the list above and write down what crosses it.
- At each boundary, ask what an attacker controls and what your agent can do next with that input.
- Rank the results by impact on your real assets. Refunds, account changes and data disclosure usually come first.
- Test the controls against an attacker who knows them, including injected content that arrives from a legitimate, authenticated peer.
One conclusion from the arXiv paper is worth keeping in view. The creation and configuration stage, where identities and trust anchors are set up, shapes every later risk, and weak choices made at onboarding are hard to fix once agents depend on each other.
Sources
- A2A Protocol Specification (sections 3.3.1, 3.6, 7, 8, 13 and 14.1.1) (accessed )
- A2A protocol definition (a2a.proto) (accessed )
- A2A documentation: Enterprise implementation of A2A (accessed )
- A2A documentation: Agent discovery (accessed )
- Anbiaee et al., Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP (arXiv 2602.11327, v2, 17 April 2026) (accessed )
- OWASP Top 10 for LLM Applications 2026 (accessed )
- OWASP Top 10 for Agentic Applications for 2026 (ASI04, ASI07, ASI08) (accessed )
- RFC 9700: Best Current Practice for OAuth 2.0 Security (accessed )
- IETF draft-ietf-webbotauth-httpsig-protocol-00, section 6.7 (SSRF) (accessed )