A voice AI vendor will quote you a per-minute rate, and the number will look
low next to what a human agent costs. What it usually will not show is that
"per minute" is stacking four separate meters — telephony, speech-to-text,
the LLM call, and text-to-speech — and that the architecture choice behind
those four meters can change the total by 10x before a single call is placed.
This is a breakdown of what a voice AI agent actually costs to run in
India in 2026: the components, the architecture decision that dominates the
bill, and the build cost on top.
The four meters running on every call
A voice agent call is billed in layers, whether or not the vendor shows them
separately:
| Component | What it does | Typical range |
|---|
| Telephony | Carries the call, inbound or outbound | ₹1–5/min |
| Speech-to-text (STT) | Transcribes caller audio | ₹1–3/min |
| LLM | Generates the response | ₹2–8/min |
| Text-to-speech (TTS) | Synthesizes the agent's voice | ₹2–6/min |
Bundled platforms roll these into one number. Reported India-facing platform
rates span roughly ₹4/min on the cheap end to ₹15–26/min once telephony,
routing, and support are stacked on top, according to a
2026 pricing breakdown from Edesy.
A platform quoting a flat low rate is usually excluding one of the four
meters, most often telephony or a routing fee — ask which ones are in the
number before comparing it to a competitor's.
The architecture choice that dominates the bill
The single biggest cost lever is not which vendor you pick — it's whether the
agent runs as a cascaded pipeline (separate STT, LLM, and TTS models
chained together) or a speech-to-speech model that processes audio
natively end to end, like OpenAI's GPT-Realtime line.
Speech-to-speech sounds like the simpler architecture, and it produces
noticeably better prosody and lower latency because it isn't round-tripping
through a text intermediate. But it is dramatically more expensive, because
the model re-processes the accumulating audio context on every turn of the
conversation rather than paying once for a short text exchange. On
Artificial Analysis's speech-to-speech pricing page,
OpenAI's GPT-Realtime-2 (High) tier lists at $1.15 per hour of input audio and
$4.61 per hour of output — multiples of what a cascaded stack pays for the
equivalent STT, LLM, and TTS calls separately.
A 2026 multi-project voice AI benchmark from DestiLabs
tracking 12 production deployments found the same split in practice:
cascaded projects landed at $0.07–$0.13 per connected minute, while
speech-to-speech projects ran $0.18–$0.21 — roughly double, with the fleet
median across all-in connected-minute cost sitting at $0.105. The latency
gap was real but smaller than the cost gap: speech-to-speech posted 540–580ms
p50 response time versus roughly 700ms p50 for cascaded, a difference most
callers on a phone line will not consciously notice.
When each one is the right call
| Cascaded pipeline | Speech-to-speech |
|---|
| Cost per minute | Lower — usually the cheapest viable option | Higher, often 1.5–10x cascaded depending on model tier |
| Latency | Good, ~700ms p50 in production | Best-in-class, ~550ms p50 |
| Debuggability | Text transcript at every step, easy to trace a bad answer | Harder — no clean text intermediate to inspect |
| Voice control | Any TTS voice, swappable independently | Locked to what the model provider offers |
| Best fit | Cost-sensitive, high-volume, scriptable flows (order status, appointment confirmation, collections) | Low-volume, high-value calls where prosody and latency matter more than unit cost |
If your unit economics need to work at call-center volume — hundreds of
thousands of minutes a month — cascaded is close to the only architecture
that clears the bar. Speech-to-speech earns its premium on calls where the
emotional register of the voice is part of the product, not just a nice-to-have.
What it costs to build, not just run
Published Indian pricing guides for voice AI builds describe three rough
tiers, per Edesy's 2026 breakdown:
- Narrow, single-intent build — around ₹75,000. One flow, scripted
fallback, no CRM integration.
- Production build with integrations — roughly ₹1.5–2 lakh. CRM or order
system connection, call logging, basic analytics.
- Custom enterprise build — ₹3.5 lakh and up. Multi-flow, multiple
languages, human handoff, compliance logging.
On top of the build, budget ongoing platform or maintenance cost separately
— reported figures run ₹5,000–10,000 a month for a production-tier
deployment, before the per-minute usage cost described above. Treat the build
quote and the running quote as two different negotiations; a low build price
with an unmodeled per-minute rate is the same trap as the "cheap now,
expensive at month three" pattern we describe for chatbots in
AI chatbot development cost in India.
Telephony: the meter that varies most by region
For India-facing traffic specifically, the telephony leg is where global and
India-focused providers diverge most. Twilio's published outbound rates to
Indian numbers run roughly ₹1.20–1.50/min, against ₹0.80–1.00/min for
India-focused carriers like Exotel, per
Edesy's Twilio-vs-Exotel comparison.
That is before counting the operational cost of a non-INR billing
relationship and the compliance work of routing calls through a carrier that
isn't natively set up for TRAI-registered numbers. For a purely India-facing
deployment, an India-focused telephony provider is usually the lower-friction
default; a global provider earns its place when the same agent also needs to
place or receive calls outside India.
The ROI comparison, done honestly
The number that makes voice AI attractive is the comparison to human agent
cost. A fully loaded human agent in a Tier-1 or Tier-2 city BPO — salary,
infrastructure, management overhead, and idle time between calls — runs
roughly ₹8–20 per minute, according to
Awaaz's 2026 call center cost breakdown.
A cascaded voice AI pipeline at ₹6–12/min all-in is genuinely cheaper on a
per-minute basis, sometimes by half or more.
The honest caveat: that comparison only holds for calls the agent actually
resolves. A voice AI agent that mishandles a call and forces the caller to
call back, or escalates to a human anyway, has spent the AI cost and the
human cost. Model the ROI against resolution rate at your actual call mix,
not against the per-minute rate in isolation — the same trap we describe for
inflated automation ROI claims generally in
what AI automation actually returns.
What to ask a voice AI vendor before you sign
- Which architecture — cascaded or speech-to-speech — and why for this
use case? A vendor who hasn't made this trade-off deliberately hasn't
priced it correctly either.
- What's included in the quoted per-minute rate? Get telephony, STT,
LLM, and TTS broken out, even if they bill it as one number internally.
- What happens to the bill at 10x call volume? Per-minute platform fees
and a self-hosted cascaded stack scale very differently.
- What's the resolution rate, measured how? "Handles the call" and
"resolves the call without a callback or escalation" are different claims.
- Is the telephony leg India-native? INR billing and a TRAI-compliant
number matter more than they look like they do at demo time.
Where to start
If you're scoping a first voice AI deployment, pick the single highest-volume
call type your team handles today — order status, appointment confirmation,
a collections reminder — and build a cascaded pipeline for that one flow
before considering speech-to-speech or a broader multi-intent agent. It is
the cheapest architecture, the easiest to debug, and it gives you a real
resolution-rate number to model the next flow against instead of a vendor's
demo.
Our AI automation page covers how we approach these builds
end to end.