Skip to content
8 min read

Voice AI for Answering Services in 2026: What It Actually Costs

Featured Image

The demo is impressive. The demo is not the job.

It was 11:47 PM. A tenant called — ceiling leaking, water spreading toward the outlet. She was scared and getting louder. The call needed someone to absorb that before anything else happened. Our operator got the maintenance coordinator on the line, gave the tenant a name and a callback time, and the situation was contained in four minutes.

That is not a demo call. That is a Tuesday.

We have been handling calls like that for 64 years. We are paying close attention to what voice AI can and cannot do in 2026 — not to dismiss it, but because we want to understand it well enough to use it right. This is what we have found.

---

How a voice AI agent is actually built

Most people think of "an AI answering service" as a single product. It is not. It is four distinct layers stacked together, and each layer introduces its own tradeoffs.

Layer 1 — Speech-to-Text (STT)

The caller speaks. The system converts audio to text before anything intelligent can happen. This sounds simple. It is not.

STT accuracy degrades with accents, background noise, fast speech, and domain-specific vocabulary. In an answering-service context that means: a property manager with a strong regional accent calling from a car, a patient from a noisy waiting room, a caller who takes 20 seconds to get to the point. Modern STT engines — Deepgram, AssemblyAI, and the models baked into voice platforms — have improved substantially. But accuracy drops are still concentrated in exactly the calls where accuracy matters most.

Layer 2 — The LLM / Reasoning Layer

Once speech becomes text, a large language model processes it and decides what to do: answer a question, route the call, take a message, trigger a workflow. This is the layer that gives voice AI its apparent intelligence.

The major model providers — OpenAI, Anthropic, Google — all offer business associate agreements for HIPAA-regulated uses, and their text-processing layers are increasingly well-covered from a compliance standpoint. But the LLM is only as good as what it was told to expect. Off-script questions, caller distress, and ambiguous situations are where the reasoning layer exposes its limits fastest.

Layer 3 — Text-to-Speech (TTS)

The system has to talk back. Modern TTS — can sound genuinely human: natural cadence, appropriate pause, emotional register that roughly fits the conversation.

"Roughly" is doing real work in that sentence. The voice quality bar has risen dramatically in the last 18 months. The bar for an answering service call — where the caller may be distressed, where how the other party sounds matters to whether they stay calm — is higher than the bar for a demo call.

Layer 4 — Orchestration and Telephony

The glue layer. This is where the call actually lands, gets routed, handed off to a human if needed, logged, and — in medical use cases — turned into a compliant message. Some platforms manage the real-time audio pipeline. Build-your-own development platforms wrap the STT/LLM/TTS layers into something a developer can configure without building the infrastructure from scratch.

This is where a lot of hidden complexity lives: latency between layers, how well human handoff works when the AI decides it cannot handle a call, how call recordings are stored and who owns them.

---

what the market looks like in mid-2026

The voice AI market for answering services has three tiers:

Dev platforms (build-your-own orchestration). These let you assemble the four layers into a custom voice agent. They are aimed at developers and companies building their own service. Pricing in this tier runs roughly $0.08–$0.10 per minute of conversation — before underlying model costs, TTS costs, and telephony costs. That headline number is real. The all-in cost is higher.

Turnkey AI answering products. Services built on top of dev platforms or natively, offering a packaged product. These typically cost more per minute than raw dev platforms but less than building custom, and handle configuration for you.

Enterprise and integrated platforms. Managed, enterprise-grade services assembled by incumbents and larger BPOs. Higher cost, more support, more compliance infrastructure.

One important distinction: major AI model providers sell business associate agreements to enterprise customers. But the BAA typically covers the text and model layer, what happens to the text after STT converts it. The audio recording layer, the actual voice file of a patient describing symptoms, may not be covered by that same BAA. More on this below.

---

The real limitations: where voice AI struggles in 2026

Voice AI handles routine calls well. It breaks on the edge cases. The problem for an answering service is that the edge cases are the job.

Latency. Even the best voice AI pipelines have detectable processing delay between the caller finishing a sentence and the bot responding. In a low-stakes transaction, confirming an appointment, this is tolerable. In a high-stakes call, a patient reporting chest tightness, a tenant in a flooding unit, a half-second delay reads as inattention. Humans fill silence with presence. Bots fill it with processing time.

Distressed callers. A caller who is scared, angry, or confused does not follow the script. They interrupt. They contradict themselves. They ask questions the system was not trained on. They need their emotional register acknowledged before the problem can even be surfaced. This is the specific skill experienced operators develop over time. It is the specific skill voice AI handles least gracefully in 2026.

Off-script and judgment calls. A caller asks something outside the configured flow. The best-designed systems hand off gracefully. Many do not. The caller waits, hears silence or a repeated prompt, and hangs up. What happens in that gap is not a user experience problem; it is a business problem.

Emergencies. An operator who hears something wrong in a caller's voice, slow speech, confusion, an answer that does not add up, can escalate before the caller asks them to. A bot processes input. It does not notice what is between the words. For medical practices and high-stakes PM scenarios, this is not an edge case. It is a Wednesday.

The "indistinguishable from human" bar. Several AI answering platforms claim their voice is indistinguishable from a human operator. A few are genuinely close in controlled conditions. In real call conditions, noisy environments, emotional calls, conversations that drift from the expected flow, the gap is still audible. And in the answering service context, audibly artificial reads as impersonal at best, unreliable at worst.

---

The HIPAA problem the demos do not mention

This matters particularly for medical practices.

The major AI platform providers have made real progress on HIPAA compliance. OpenAI offers a business associate agreement for enterprise customers. Anthropic covers its API under BAA for compliant uses. Retell AI offers a self-service BAA on all plans — a notable move for a dev platform.

But as of mid-2026, there is a documented gap: the audio layer.

When an AI voice agent handles a patient call, the audio is being processed in real time, a recording may be retained, the STT conversion produces a text record, and the LLM processes that text. The BAA coverage on major platforms, including OpenAI's Realtime API, as of mid-2026, covers the text and model layer. The audio recording itself, the actual sound file of the patient describing symptoms, is often not explicitly covered by that same BAA.

This is not speculation, but rather the current limitation of how the compliance infrastructure has been built. It may be changing, but "changing" is not the same as "changed." If your practice uses an AI voice service and you ask the vendor whether their BAA covers audio recordings specifically, listen carefully to the answer. A vague answer is not a compliant answer.

Human operators working under a properly executed BAA, one that covers the entire call handling workflow including any recording, remain the cleaner compliance path for medical voice use today.

---

What it actually costs

The per-minute headline number obscures several layers.

The raw platform cost. Dev platforms like can run roughly $0.08–$0.10/minute. For simple, scripted call types, this is materially lower than live human operators.

The underlying model and TTS costs. The dev platform fee sits on top of separately metered API costs for the LLM and TTS layers. The all-in cost per minute for a real configured deployment is higher than the platform headline — a rough estimate puts it in the $0.15–$0.25/minute range depending on model tier and voice quality settings, before accounting for any compliance add-ons.

HIPAA compliance mode. Platforms offering HIPAA-compliant configurations often charge a premium, a "compliance mode" or "HIPAA plan" tier. The base per-minute pricing typically does not include it. Budget for this explicitly.

Build and configuration cost. A dev platform is not a product you turn on. It requires a developer to configure the agent, design call flows, train it on your specific scenarios, and maintain it as call patterns evolve. That labor cost is real and ongoing.

The cost of the calls AI does not handle well. This is the number that does not appear on the pricing page. Every distressed patient who hangs up, every property emergency that does not get properly escalated, every new prospect who gets a flat answer and calls the next number on the list — each has a cost. It does not show up in the per-minute dashboard. It shows up in retention data and review volume over the following year.

Additionally, model drift is a real thing. These tools are not set and forget. If the workflow breaks it does not necessarily let you know. It requires supervision and maintence.

For high-volume, highly scripted, low-stakes call types, appointment confirmations, basic information requests, hours and location queries, the economics of AI are genuinely attractive. For the calls that matter most in medical, PM, and trades contexts, the cost of a failed interaction is not captured in a price-per-minute comparison.

---

Five questions to ask any voice AI vendor before you buy

Whether you are evaluating AI for your own operation or vetting a service that uses AI in its workflow, these are the questions that matter:

  1. What call types can your system not handle? An honest vendor has a clear, specific answer. Evasion is information.
  1. Does your BAA cover audio recordings and voice data, or only text? This is the HIPAA gap question. A BAA that covers the text transcript but not the underlying audio leaves a material portion of patient data outside your safeguards.
  1. How does handoff to a human work — and how long does it take? Ask to hear a recording of a handoff that did not go smoothly. Every platform has them.
  1. What happens when the system does not understand a question? Silence, repetition, or graceful handoff — the answer tells you a lot about how the system was designed.
  1. What does your dashboard not show? Call completion rates look clean; caller experience does not always match. Ask how they measure calls that completed but lost the customer anyway.

---

Where we sit on this

We are paying close attention to what voice AI can do, because we think the right answer in this space will involve understanding the technology well enough to use it where it genuinely helps and keep people where it genuinely matters. We are evaluating that question seriously, and piloting several different offerings, and we do not have a finished answer yet.

What we do have is 64 years of knowing what the hard calls sound like. We are not convinced any current platform clears the bar for the calls where that judgment is the job. We are watching closely.

If you are evaluating voice AI for your medical practice, PM portfolio, or trades operation — or if you want to talk through what your current after-hours coverage actually handles — we are happy to have that conversation.

Call us at 800.248.2255 or reach out at support@amessagecenter.com.

People Answering People. Since 1962.