Bad call quality starts in one of seven VoIP infrastructure layers. See what each one does and where calls break.

Takeaways
VoIP infrastructure is every layer a call crosses between the handset and the public phone network (PSTN), so a choppy, one-way, or dropped call can be traced to the specific layer that caused it.
VoIP infrastructure is the full chain of systems that carries a voice call over IP and hands it to the phone network.
Definition: VoIP infrastructure covers every device, network, and service a call crosses from the caller's handset to the called party, including the session border controller, the SIP trunk, and the PSTN interconnect that sit outside the office.

The common picture stops at the office: enough bandwidth, QoS on the router, and the right phones. That covers two layers out of seven. The session border controller (SBC), the SIP trunk, and the PSTN interconnect run outside the building, and they decide call quality just as often. Because the infrastructure is a chain, every quality problem has a location in it.
A VoIP call moves through four steps. First, the phone or app turns speech into packets using a codec such as G.711. Second, SIP signalling sets the call up: it finds the other party, agrees on a codec, and tells each side where to send audio. Third, the audio itself travels separately as RTP packets, usually every 20 ms. Fourth, when the other party is on a normal phone number, a carrier converts the call and hands it to the PSTN.
Signalling and media take separate paths, so a call can connect cleanly and still have no audio. That split explains most of the faults covered below.

Connect your phone system to a carrier network with Telnyx SIP trunking.
A call crosses seven layers in a fixed order, and knowing which layer a symptom belongs to is how you find the fault instead of guessing. The first two layers are yours. The last five usually belong to a provider or a carrier, which is why a clean office network does not guarantee a clean call.
What is wrong with your call?
Pick the symptom you hear to get the likely cause and the fix.
A private IP address is still in the SDP
Capture the INVITE and the 200 OK and read the c= line in the SDP. If it shows an address in 10.x, 172.16.x to 172.31.x, or 192.168.x, the far end is sending RTP to an address it can never reach.
Turn off SIP ALG on the router first, because it often rewrites the headers but not the SDP. Then point the phones or PBX at an SBC with NAT traversal enabled, so it learns the real public address and port from the first RTP packet it receives.
Before: c=IN IP4 192.168.1.20 in the 200 OK, and the caller hears silence. After: the SBC rewrites it to c=IN IP4 203.0.113.10, and audio flows both ways.
The ACK never reached the answering side
SIP gives an answered call 64 times the 500 ms retransmit timer, which is 32 seconds, to receive an ACK. Without it, the answering side keeps resending the 200 OK, gives up, and sends BYE.
The ACK is usually sent to a private address taken from the Contact header. Fix the Contact header through the SBC or NAT settings so it carries the public address, and the drop at 32 seconds disappears.
Before: the 200 OK is resent about ten times and the call ends at 0:32. After: the ACK arrives within one round trip and the call stays up for its full length.
The upload is full or voice is unmarked
Multiply your peak concurrent calls by 80 kbps for G.711 and compare that with your upload speed, not your download. Most office links are asymmetric, so the upload runs out first.
Mark RTP as DSCP EF (46) and SIP as CS3 (24), and give EF strict priority on the WAN interface. If the link is still too small, G.729 uses about 24 kbps per direction with IP, UDP, and RTP headers at 20 ms packets.
Before: 25 G.711 calls need 2 Mbps up on a 2 Mbps upload, and audio chops at 10 am. After: the same 25 calls on G.729 need about 600 kbps.
Jitter or packet loss is beating the buffer
Each RTP packet carries 20 ms of speech, so every lost packet is a 20 ms hole. G.711 without concealment starts to sound broken at around 1% loss. Check the loss and jitter figures in the RTCP reports or in your phone's call statistics.
Move softphones and IP phones off Wi-Fi onto a wired port first. Raise the adaptive jitter buffer only if loss is already low, since every extra 20 ms of buffer is taken out of the 150 ms delay budget.
Before: 3% loss and 60 ms jitter on a laptop on Wi-Fi. After: the same laptop on Ethernet shows under 0.5% loss and under 20 ms jitter.
One-way delay is past 150 ms
Ping the carrier's SBC and halve the round trip to get the network share. Add about 20 ms for the codec packet and the size of the jitter buffer. If the total passes 150 ms, callers start to interrupt each other.
The quickest gain is usually distance. Register to the carrier point of presence nearest your users, or use private transport, instead of routing every call through one distant region.
Before: 120 ms network plus 60 ms jitter buffer is 180 ms one way. After: a closer region at 40 ms plus the same buffer is 100 ms.
The table maps each layer to its usual owner and the sound of its failure.
The seven layers of VoIP infrastructure
| Layer | Who usually runs it | What you hear when it fails |
|---|---|---|
| 1. Endpoints (IP phones, softphones, browsers) | You | Echo, muffled audio, or a phone that will not register |
| 2. Local network (LAN, Wi-Fi, router, firewall) | You | Choppy audio at busy hours, one-way audio |
| 3. Internet or private transport | Your ISP or your provider | Delay, robotic or clipped speech |
| 4. Session border controller | You, your provider, or both | One-way audio, calls that drop at a fixed time |
| 5. SIP trunk or carrier network | Your provider or a carrier | Failed call setup, poor quality on certain routes |
| 6. PSTN interconnect | A carrier | Calls that fail or degrade only to certain countries or networks |
| 7. Phone numbers | Your provider or a carrier | Outbound calls labeled as spam or left unanswered |

Layers 3 to 7 are where call quality is decided outside your control, and they are the layers a carrier tunes. James Whedbee, who works on the network Telnyx operates, names the controls that sit there:
"Telnyx made the infrastructure part of the product. That gives us a different set of levers: network routing, media anchoring, SIP behavior, number reputation, failover, and AI placement." James Whedbee, VP of Engineering at Telnyx |
Each lever acts on a layer from the table:
Endpoints turn speech into packets and play the far end back. When they fail, you hear echo from poor cancellation, a one-sided call from a bad headset, or a phone that cannot register. The local network carries those packets to the edge. Congestion, Wi-Fi interference, and a firewall that blocks RTP ports show up here, and this is the layer the standard setup checklist already covers.
Transport is the internet link or private circuit between your site and your provider. The SBC sits at the edge of a network, cleans up addressing, and relays media. The SIP trunk connects your phone system to a carrier network so calls can reach real phone numbers. If you are deciding between a trunk and a hosted service, the trade-offs are covered in SIP trunking vs VoIP. Faults here tend to look like network problems in the office, which is why they get misdiagnosed.
The PSTN interconnect is where a carrier exchanges calls with other phone networks. A weak route here explains calls that fail or sound poor only to some countries or carriers. Phone numbers sit at the end of the chain. Teams that get a VoIP number also inherit how that number is treated by other networks, which affects answer rates on outbound calls. The provider that issues your phone numbers controls porting, registration, and how those numbers are presented.
A VoIP call needs low one-way delay, steady packet timing, very little loss, and a predictable amount of bandwidth per call. Two of those targets rest on a standard or on arithmetic. The other two depend on the codec and the jitter buffer, so no single published threshold applies:
Per-call bandwidth comes from three numbers: the codec rate, the packet interval, and the header size. Redo the math for your own codec and call volume:
The 100 kbps per call figure that circulates is this 80 kbps plus Ethernet framing and some headroom. Wideband codecs such as G.722 improve clarity, and Opus adjusts its bit rate to the network, so check the rate your phones actually negotiate before sizing the link.
Delay accumulates in a set order, and each stage spends part of the 150 ms budget:
The first three stages are mostly fixed. Network transit is the stage that grows with distance, and it is the one the routing and media anchoring levers act on.
Why echo shows up on IP calls: people rarely notice echo while the round trip stays very short, commonly put at under about 50 ms. IP paths often run longer than that, so every endpoint needs working echo cancellation, and an echo complaint is often a delay problem rather than a faulty handset.
QoS works at two layers of the local network. At Layer 2, 802.1p tags inside 802.1Q frames tell switches to forward voice first. At Layer 3, DSCP markings (Expedited Forwarding, value 46, is the usual choice for voice) tell routers the same thing. A dedicated voice VLAN keeps phones apart from bulk data traffic.
These markings only hold on networks you control. Once packets reach the public internet, most networks ignore them. That is the point where a call that tested fine in the office can still degrade, and where layers 3 to 5 take over.
NAT breaks calls because it rewrites the address in a packet's header but not the address SIP carries inside the message body. Under RFC 3261, a SIP message includes a session description (SDP) that states where the other side should send audio. A phone behind a router writes its private address into that SDP. The router rewrites the outer header to its public address and leaves the SDP alone. The far end then sends audio to a private address it cannot reach, and you get one-way audio or none.
Many routers ship with SIP ALG, a feature that tries to fix this by rewriting the addresses inside SIP messages. In practice it often corrupts the messages instead, and switching SIP ALG off is a standard first step when troubleshooting.
Use the symptom to find the layer:
A session border controller sits at the edge between two networks and handles the jobs NAT cannot:
Securing VoIP infrastructure means encrypting both signalling and media, blocking fraud at the trunk, and meeting the rules that apply to emergency calls and caller ID.
Signalling and media need separate protection. TLS encrypts SIP signalling. SRTP, defined in RFC 3711, encrypts and authenticates the RTP audio using AES. Older guidance that lists DES or RC4 for voice describes ciphers that are no longer considered secure. When a vendor page says "end-to-end encrypted," check whether it names both TLS and SRTP.
The four threats every VoIP deployment faces map to specific controls:
Emergency calling is part of the infrastructure. Each user's dispatchable location has to be registered and kept current so a 911 call reaches the right dispatcher. The details are covered in E911 requirements.
Caller ID authentication is the second regulatory layer. STIR/SHAKEN is the framework US carriers use to sign and verify caller ID, and it affects whether your outbound calls display as verified.
Ask any provider three questions:
An uptime SLA is a downtime allowance, and it says nothing about how a provider fails over.
A year has 525,600 minutes. Each extra nine cuts the allowance by a factor of ten.
The percentage leaves out what matters during an outage. Ask whether calls in progress survive a failure, whether the service runs in more than one region, and how the SLA is measured and credited.
Some redundancy is yours to build before go-live. Work through it in order of defense:
You can build VoIP infrastructure in three ways, and the real difference between them is which of the seven layers you run and which you hand over.
Self-hosted. Open-source software such as Asterisk or FreeSWITCH runs on your own servers, with a SIP trunk for PSTN access. You run layers 1 to 4 and pay in hardware and staff time. It suits teams that need deep customization and have voice engineers on hand.
Hosted phone system. A provider runs the call platform and you buy seats. You run layers 1 and 2, and pricing scales per user. It suits offices that want a working phone system with little engineering.
Carrier SIP trunking or voice API. You build on a carrier's network through SIP trunks or a VoIP API, and pricing scales with usage. You choose how much of layers 1 to 4 to run, and the carrier runs layers 5 to 7. Telnyx is an example of this model: it operates carrier infrastructure, call routing, media handling, phone numbers, and SIP itself.
Ownership affects price as well as quality. A provider that resells another carrier's network passes along that carrier's price changes and cannot fix that carrier's routes.
"Having a whole stack ownership lets us also control the pricing of it, so that we can charge a fair amount and not increase the charges overnight just because one of our vendors has decided to increase the price." Abhishek Sharma, Senior Technical Marketing Manager at Telnyx |
A rollout goes in four phases:
An AI voice agent puts speech recognition, a language model, and speech synthesis into the call path, so every millisecond the network spends comes out of the time the agent has to respond. A route that adds delay a person would barely notice can make an agent sound slow or cause it to talk over the caller.
That makes AI call quality an infrastructure question as much as a model question:
"We are not only orchestrating an AI agent. We also operate carrier infrastructure, call routing, media handling, phone numbers, SIP, and the voice AI layer. That matters because production voice quality is not only a prompt or model problem. It is also routing, answer rates, codec handling, failover, observability, and operational accountability when something goes wrong." James Whedbee, VP of Engineering at Telnyx |
Routing and media anchoring set how far audio travels. Answer rates depend on number reputation. Codec handling decides what the speech recognizer actually hears. Failover and observability decide whether a problem is caught and fixed. The measured results of running the agent next to the call path are in the write-up on co-located infrastructure.
Before changing the model: if an AI agent sounds slow or drops calls, check the seven layers in the architecture table above before rewriting the prompt or swapping the model.
What is VoIP and how does it work?
VoIP (voice over IP) carries phone calls as data packets instead of over copper lines. A codec turns speech into packets, SIP signalling sets up the call, and RTP carries the audio. When the other party is on a regular phone number, a carrier hands the call to the public phone network so it can ring any phone.
What are the minimum requirements for VoIP infrastructure?
Plan for about 80 kbps per direction for each G.711 call before Layer 2 overhead, one-way delay at or under 150 ms, and low, steady jitter and packet loss. Add QoS or a voice VLAN on your network, a firewall that allows SIP and RTP, and a SIP trunk or hosted service to reach the phone network.
What causes poor call quality after a VoIP installation?
The usual causes are congestion on the local network, NAT or SIP ALG problems at the router, too much delay or loss on the internet path, and weak carrier routes to certain destinations. Match the symptom to a layer: one-way audio points to NAT, choppy audio at busy hours points to the local network, and problems on certain routes point to the carrier.
What factors contribute to VoIP latency and delay?
Delay builds up from codec processing, packetization (usually 20 ms per packet), the jitter buffer, network transit, and the hop into the public phone network. The first three are mostly fixed. Network transit grows with distance, so routing and where the audio is relayed have the largest effect on how much delay a call collects.
Can I keep my existing phone numbers when switching to VoIP?
In most cases, yes. Number porting moves your existing numbers from the old carrier to the new provider. It can take days or weeks depending on the carrier and country, so start porting early in the rollout and keep the old service active until every number has moved and been tested.
Every bad call has a layer. Telnyx operates the carrier layers itself, including call routing, media handling, phone numbers, and SIP, so you can trace a problem past your office network.
Explore SIP trunkingRelated articles
AI-to-human handoff for voice AI agents: a practical guide
How Telnyx Voice AI Handles Multiple Speakers on One Call
How does AI voice work? The audio path from caller to reply
The 9 best WhatsApp API providers in 2026

Finding the Right Indian Text-to-Speech Voice for Your Application
%20(1).png?width=96&format=webp)
Model serving: how to run open-weight LLMs close to your users
