Voice AI vs A2A for customer service: serving callers and serving agents
Voice AI agents answer people on the phone in natural speech. A2A accepts structured tasks from customers' agents. How they differ and share a back end.
A voice AI agent is for serving people who phone a business: it listens to natural speech, works out what the caller wants, acts on the business’s systems, and answers in a synthesized voice. A2A (Agent2Agent) is for serving software agents that contact the business on a customer’s behalf: they send structured tasks over HTTPS to an endpoint described by an Agent Card.
Both can sit in front of the same service logic. What differs is the caller. Voice AI replaces the menu tree with conversation for people. A2A gives customers’ agents a route that never turns data into speech in the first place.
Status as of September 26, 2026. Voice AI agents are vendor products and platforms. They reach the phone network through telephony providers and SIP (RFC 3261), and there is no common standard for the agent layer. The FCC confirmed in Declaratory Ruling FCC 24-17, released February 8, 2024, that AI-generated voices count as “artificial” voices under the Telephone Consumer Protection Act. A2A is a Linux Foundation project at protocol version 1.0; its latest release, v1.0.1, shipped on May 28, 2026.
What voice AI agents are for
A voice agent handles a live call in a loop: speech in, text or audio into a model, an action or answer out, speech back. Vendor documentation shows the plumbing:
- Twilio ConversationRelay. A call is connected with the TwiML
<Connect><ConversationRelay>instruction. Twilio handles speech recognition and text-to-speech, sends the caller’s words as text to the developer’s application over a WebSocket, and speaks the application’s replies back to the caller. - OpenAI Realtime API over SIP. A SIP trunk sends calls to OpenAI’s SIP endpoint. An incoming call triggers a
realtime.call.incomingwebhook, the application accepts or rejects it, and a Realtime session handles the speech on the call.
Done well, this serves people better than a keypad menu. Callers say what they want in their own words, can interrupt, and get help at any hour. The design target is a human caller: turn-taking, tone, clarifying questions, and the handoff to staff when the model reaches its limits.
What A2A is for
A2A assumes the caller is software. The business publishes an Agent Card, usually at /.well-known/agent-card.json, with its skills, endpoints and required authentication. A customer’s agent sends a message with text and data parts. The business’s agent creates a task that can complete, fail, be refused, or pause for more input or for authorization, and returns results as artifacts. Tasks can outlive a connection, and the client can poll, stream or receive push notifications.
The same request both ways
Illustrative: a customer wants appointment APT-5521 moved to Thursday afternoon.
On a call answered by a voice agent, every value makes a round trip through audio:
1. Caller speaks: "I need to move my appointment, A-P-T five five two one, to Thursday afternoon."
2. Speech recognition transcribes it, possibly as "APT 5-5-2-1" or "a PT 5521".
3. The model reads the transcript, confirms the ID, and checks available slots.
4. Text-to-speech reads options aloud; the caller picks one; the model confirms.
5. The call ends. The business keeps its own recording and notes.
If the caller is itself an AI agent, steps 1 and 4 are synthesized speech on both ends of the line.
Over A2A, the customer’s agent sends the request as data. Illustrative, JSON-RPC binding:
{
"jsonrpc": "2.0",
"id": 12,
"method": "SendMessage",
"params": {
"message": {
"messageId": "2f6c8e0a-91d4-4b7e-a3f5-6d2c1b0e9a87",
"role": "ROLE_USER",
"parts": [
{ "text": "Please move this appointment to Thursday afternoon." },
{
"data": { "appointmentId": "APT-5521", "preferredDate": "2026-10-01", "preferredWindow": "12:00-17:00" },
"mediaType": "application/json"
}
]
}
}
}
The business’s agent can reply with a completed task and the new slot as an artifact, or with TASK_STATE_INPUT_REQUIRED and a list of open times to choose from. Nothing is transcribed, and both sides hold the same task ID and result.
Side by side
| Voice AI agent | A2A | |
|---|---|---|
| Purpose | Serve people who call, in natural speech | Accept structured tasks from customers’ agents |
| Layer | Speech application on the telephone network | Application protocol over HTTPS or gRPC |
| Who talks to whom | A caller (usually a person) and the business’s voice agent | A client agent and the business’s remote agent |
| Transport | Phone network or SIP audio; vendor bridges pass audio or text to the model | JSON-RPC 2.0, gRPC or HTTP+JSON; streaming and webhook push |
| Discovery | A phone number | Agent Card at /.well-known/agent-card.json, registries or direct configuration |
| Auth | Caller ID, attested for the number by STIR/SHAKEN, plus in-call verification questions | Security schemes declared in the Agent Card; in-task authorization through TASK_STATE_AUTH_REQUIRED |
| State | The call; recordings and notes kept by each side | Tasks that outlive any connection; contextId groups related work |
| Governance and status | Vendor platforms; SIP (RFC 3261); FCC rules on artificial voices apply to certain calls | Linux Foundation project; version 1.0 (v1.0.1, May 28, 2026) |
When to use each
Use voice AI for people who call. It can answer at any hour, understand free-form requests, and route the hard cases to staff. It is the right upgrade for a phone line that people use.
Use A2A for callers that are software. When a customer’s agent wants an order status, a refund or a new appointment time, a structured task is exchanged in one request instead of a spoken conversation, the business gets a place to authenticate the caller, and both sides keep a record of the same task.
Using both
Run them on one back end. The voice agent and the A2A agent should call the same systems and apply the same policies, so a customer gets the same answer on either channel. If you already run an AI service agent for voice or chat, putting an A2A entrance in front of it reuses that work. Emissar’s Front Door is a hosted A2A endpoint in front of a business’s existing service agent or systems; its status is Open to design partners.
The two channels also meet at escalation. A voice agent hands hard calls to staff. An A2A task can pause in TASK_STATE_INPUT_REQUIRED while a person on the business’s side reviews it, then continue on the same taskId. When two voice agents meet on a phone call looks at the case where both ends of a call are machines.
Common misconceptions
“A good voice agent makes A2A unnecessary.” A voice agent improves the call for people. When the caller is software, both sides still turn structured data into speech and back, and neither side can verify whom the other represents.
“A2A is a voice protocol.” A2A carries text, files and structured data. The specification lists audio and video among the content it can exchange, as file references, but it does not place or carry phone calls.
“Caller ID tells a voice agent who is on the line.” STIR/SHAKEN lets carriers vouch for the calling number. It says nothing about which software placed the call or which customer it represents.
“Voice and A2A need separate policies.” They need separate front ends and one set of rules. Different answers on different channels invite customers, and their agents, to shop for the better one.
Questions
- Can a voice AI agent and an A2A agent share the same logic?
- Yes. Both can call the same order, booking and billing systems and apply the same policies. The voice agent adds speech recognition, speech synthesis and turn-taking for people. The A2A side adds an Agent Card, caller authentication and task handling for software.
- Do FCC rules on AI voices apply to agent-to-agent phone calls?
- The FCC ruled in February 2024 (FCC 24-17) that AI-generated voices are artificial voices under the Telephone Consumer Protection Act, which restricts certain calls that use them. Whether a particular call falls under those rules depends on the call. A2A requests are not phone calls at all.
Sources
- Twilio Docs: ConversationRelay (updated August 24, 2026) (accessed )
- OpenAI API documentation: Telephony and SIP (accessed )
- RFC 3261: SIP: Session Initiation Protocol (June 2002) (accessed )
- Combating Spoofed Robocalls with Caller ID Authentication (FCC) (accessed )
- FCC Confirms that TCPA Applies to AI Technologies that Generate Human Voices (Declaratory Ruling FCC 24-17, CG Docket 23-362, released February 8, 2024) (accessed )
- A2A Protocol Specification (sections 1.2, 3, 4, 7.6 and 8) (accessed )
- A2A releases on GitHub (v1.0.1, May 28, 2026) (accessed )
- Emissar Front Door (module page and status) (accessed )