The general build-versus-buy framework for chatbots applies to voice as a starting point, but telephony adds real-time infrastructure, connectivity, concurrency and continuous-uptime demands that chat never had to deal with — and underestimating them is where custom voice builds go over budget.
The same question, a harder version
The general build-versus-buy framework for chatbots applies to voice as a starting point — commodity versus specific, hybrid rather than binary, team capability as a real constraint. But telephony adds a layer of infrastructure that chat never had to deal with, and it changes the calculation in ways worth treating separately rather than assuming voice is simply a chatbot with audio attached.
Building a voice AI system from scratch means building or integrating speech recognition, speech synthesis, telephony connectivity, real-time low-latency infrastructure, and failover — on top of everything a chatbot already needs. Each of these is its own substantial engineering domain, and underestimating any one of them is where custom voice builds most often go over budget and over schedule.
What telephony specifically adds to the calculation
Five domains that do not exist in a text-based build at all, and that a platform absorbs as part of its offering.
-
1
Real-time infrastructure
Voice cannot tolerate the latency that is invisible in a chat interface. Building infrastructure that reliably delivers low-latency streaming recognition and synthesis, at scale, under real network conditions, is a specialised and continuing engineering discipline, not a one-time integration task.
-
2
Telephony connectivity itself
SIP trunking, PBX integration, number provisioning — this is an entire domain most software teams have no existing expertise in, and it typically requires an ongoing relationship with a telephony provider regardless of which conversational technology sits behind it.
-
3
Concurrent call capacity and scaling
Voice infrastructure has to handle real-time concurrent sessions in a way that scales predictably, which is architecturally different from scaling a typical web or chat backend and needs its own capacity planning discipline.
-
4
Uptime and failover expectations
Phone systems are expected to work continuously, with defined behaviour when something fails rather than a generic error page. Building this reliability from scratch is considerably more demanding than for a typical web application.
-
5
Recording, compliance and telephony-specific data handling
Call recording, retention and consent requirements are specific to voice and layer on top of the general data handling any AI system needs.
Where the commodity-versus-specific line usually falls for voice
Applying the same component-by-component thinking as the general framework, with voice-specific weight on where things land.
- Telephony connectivity and infrastructure: almost always commodity. Building this yourself rarely produces a better outcome than a platform that has already solved it at scale, and the ongoing operational burden of running it yourself is substantial.
- Speech recognition and synthesis: almost always commodity, for the same reason — this is a deep, continuously advancing specialisation that a platform amortises across many customers.
- Conversation logic and knowledge: the same commodity-versus-specific judgement as a chatbot — standard qualification and FAQ patterns are commodity, unusual workflow-specific logic may not be.
- Deep system integration: the same judgement as a chatbot again — routine API reads are commodity, integration with unusual internal or legacy systems may need custom work regardless of the conversational layer chosen.
- Compliance and recording requirements specific to your regulatory environment: sometimes genuinely specific enough that no available platform meets them exactly, which points toward a hybrid where the platform handles conversation and a custom layer handles compliance-specific recording and retention.
The pattern that emerges is stronger toward buying for voice specifically than for chat, because the infrastructure layer is both more specialised and more continuously demanding to maintain than anything equivalent in a text-only system.
What building voice AI from scratch actually requires
Being realistic about the scope, for organisations genuinely considering this path.
- A telephony integration relationship and the ongoing expertise to manage it — this does not end at initial setup and needs continuing attention as providers and requirements change.
- Real-time infrastructure engineering expertise, which is a different discipline from typical backend or web engineering and is not a skill most software teams already have in-house.
- Ongoing monitoring and incident response specifically for a system expected to run continuously, with materially higher availability expectations than most internal tools.
- Speech recognition and synthesis capability that keeps pace with a fast-moving field, or acceptance of maintaining an increasingly dated approach relative to what platforms are advancing toward.
- A failover and resilience design covering model, network, telephony and tool failures, tested regularly rather than assumed to work.
- All of the general chatbot build requirements as well — knowledge maintenance, testing infrastructure, guardrail enforcement — layered on top of everything above.
This is a substantially larger undertaking than building a chatbot, and it should be sized accordingly during any build-versus-buy evaluation — treating voice as chat-plus-a-small-audio-layer is the most common and most costly misjudgement in this decision.
When building voice AI genuinely makes sense
Narrower than for chat, but real in specific circumstances.
- Your organisation already has genuine telephony and real-time systems expertise — a telecom company, for instance — making this closer to adjacent work than to an entirely new discipline.
- Voice interaction is a core product differentiator for your business, not an operational support function, justifying the investment in owning the full stack.
- Regulatory or data residency requirements specific to your situation cannot be met by any available platform, confirmed through direct evaluation rather than assumption.
- You are operating at a scale where the ongoing cost of a platform's pricing model genuinely exceeds the fully loaded cost of building and maintaining your own infrastructure — modelled explicitly, not estimated.
- You need call handling logic so specific to an unusual internal process that no platform's configuration can accommodate it, confirmed by testing against real platforms rather than assumed in advance.
Cost comparison specific to voice
The same discipline as the general framework, with voice-specific line items that are easy to omit from an initial build estimate.
- Telephony connectivity costs, both the provider relationship and the ongoing integration maintenance, are a distinct and continuing line item separate from any conversational AI cost.
- Real-time infrastructure has its own scaling cost curve, different from a typical web application, and needs its own capacity planning rather than reusing assumptions from other systems.
- Compliance and recording infrastructure specific to voice is a real, ongoing cost in a build estimate and is frequently underestimated or omitted entirely from initial projections.
- Model platform pricing at your actual expected call volume, including peaks, not average volume — voice traffic patterns are often peakier than chat traffic and pricing models sensitive to concurrency can behave differently under real conditions than a simple average suggests.
- Include the cost of the specialised engineering talent a build requires — real-time and telephony expertise commands its own market rate and is not interchangeable with general software engineering capacity.
Model concurrent capacity needs specifically, not just total call volume — pricing models sensitive to concurrency behave very differently under real peak conditions than an average-volume estimate suggests, and this is worth modelling explicitly before comparing build and buy costs.
Decision framework and next step
Four questions specific to the voice decision, on top of the general framework.
-
1
Does your organisation have genuine existing telephony and real-time systems expertise?
If not, that gap alone is a strong argument toward buying the infrastructure layer, whatever the answer looks like for conversation logic.
-
2
Have you sized the telephony and real-time infrastructure requirement realistically, separate from the conversational AI requirement?
Treating voice as chat-plus-audio is the most common and costly misjudgement in this decision.
-
3
Have you confirmed platform limitations through direct testing, not assumption?
Before concluding no platform meets a specific requirement, test that conclusion against current platform capability rather than an earlier or assumed evaluation.
-
4
Have you modelled cost at peak concurrent capacity, not average call volume?
Voice traffic patterns and concurrency-sensitive pricing behave differently under real peaks than an average-volume estimate suggests.
For most organisations, the telephony and infrastructure layer should be bought regardless of how conversation logic and integration depth are decided — the general commodity-versus-specific framework still applies to those layers. Where buying, the buying checklist covers what to evaluate; our custom software services page covers the build side where it is genuinely warranted; our AI solutions overview covers staging.
Frequently asked questions
-
1
How is voice AI build-versus-buy different from chatbot build-versus-buy?
Telephony adds real-time infrastructure, SIP and PBX connectivity, concurrent call capacity planning, continuous-uptime expectations with defined failover, and voice-specific recording compliance — none of which exist in a text-based build. These layers push the decision more strongly toward buying than the equivalent chatbot decision.
-
2
What does building voice AI from scratch actually require?
A telephony integration relationship and the expertise to manage it, real-time infrastructure engineering — a distinct discipline from typical backend work — continuous monitoring and incident response, and a tested failover design, layered on top of everything a chatbot build already needs.
-
3
Where does the commodity-versus-specific line usually fall for voice?
Telephony connectivity, recognition and synthesis are almost always commodity — best bought from a platform that has solved them at scale. Conversation logic and deep system integration follow the same case-by-case judgement as a chatbot.
-
4
When does building voice AI genuinely make sense?
When the organisation already has genuine telephony and real-time systems expertise, when voice is a core product differentiator rather than an operational function, when specific regulatory requirements cannot be met by any available platform, or at a scale where the fully loaded build cost genuinely undercuts platform pricing — each confirmed by testing, not assumed.
-
5
What is the most common costly mistake in this decision?
Treating voice as a chatbot with an audio layer attached, rather than sizing the telephony and real-time infrastructure requirement as its own substantial and continuing engineering domain.
For most organisations the telephony and real-time infrastructure layer should be bought, whatever the answer turns out to be for conversation logic and integration depth.