You look at your SIP trunk provider’s dashboard, see a clean 64 kbps figure for G.711 calls, and assume you have plenty of headroom. Then the network chokes during peak hours, calls drop, and voices sound robotic. The culprit isn’t the audio itself-it’s the invisible tax every packet pays to travel across your network. This is the gap between codec payload and true per-call bandwidth.
If you are planning capacity for a new office or troubleshooting quality issues in an existing deployment, relying on nominal codec rates is a trap. A G.711 call doesn’t just send voice; it sends headers, framing bits, and encryption wrappers that can double or even triple the actual bandwidth consumption. Understanding this distinction saves you from over-provisioning expensive circuits or under-provisioning critical links.
The Anatomy of a VoIP Packet
To grasp why bandwidth numbers don’t match reality, you need to peel back the layers of a single voice packet. Imagine a digital envelope. The letter inside is the audio data (the payload). But you can’t just throw the letter into the wind; you need an address, a stamp, and a tracking number. In networking terms, these are your protocol headers.
Every VoIP packet travels through a stack of protocols, each adding its own fixed size. On a standard Ethernet LAN with IPv4, the breakdown looks like this:
- RTP Header: 12 bytes. It tracks sequence numbers and timestamps to keep audio in order.
- UDP Header: 8 bytes. It handles port numbers so the receiving device knows which application gets the data.
- IPv4 Header: 20 bytes. It provides the source and destination IP addresses.
- Ethernet II Frame: 18 bytes. This includes the preamble, start delimiter, MAC addresses, and the frame check sequence.
Add those up, and you get 58 bytes of overhead per packet. Notice something? That number doesn’t change based on how much audio you’re sending. Whether you’re using a high-fidelity codec or a low-bitrate one, you pay this 58-byte tax on every single packet sent. If you use VLAN tagging, add another 4 bytes. If you use SRTP for encryption, the payload might grow slightly, but the header structure remains the primary driver of inefficiency.
Why Packetization Interval Matters More Than You Think
Here is where most administrators get tripped up: the packetization interval (ptime). This is how much audio you bundle into each packet before sending it. The industry standard is often 20 milliseconds (ms), meaning you send 50 packets per second (pps).
Let’s do the math for a G.711 call. G.711 has a payload rate of 64 kbps. At 20 ms, that’s 160 bytes of audio per packet. Add the 58 bytes of overhead, and each packet is 218 bytes total.
Now, calculate the bandwidth:
- 218 bytes/packet × 50 packets/second = 10,900 bytes/second.
- 10,900 bytes × 8 bits/byte = 87,200 bps (87.2 kbps).
See the difference? The codec says 64 kbps. The network sees 87.2 kbps. That’s a 36% increase purely from overhead. Now, imagine you change the ptime to 30 ms. You now send fewer packets (approx. 33 pps), but each packet carries more audio. The overhead cost drops because you’re paying the 58-byte tax less frequently. However, increasing ptime adds latency. If you go too high (e.g., 40 ms+), you might save bandwidth but introduce noticeable delays in conversation.
Real-World Bandwidth Figures by Codec
Different codecs handle compression differently, which changes how painful the overhead tax feels. For uncompressed codecs like G.711, the overhead is significant but manageable. For highly compressed codecs like G.729, the overhead becomes the dominant factor.
| Codec | Nominal Payload Rate | Payload Size (20ms) | Overhead (Bytes) | Total One-Way Bandwidth | Recommended Provisioning |
|---|---|---|---|---|---|
| G.711 | 64 kbps | 160 bytes | 58 | ~87.2 kbps | 100 kbps |
| G.729 | 8 kbps | 20 bytes | 58 | ~31.2 kbps | 40 kbps |
| Opus (Standard) | ~24-48 kbps | Variable | 58 | ~48-72 kbps | 80 kbps |
| G.722 | 64 kbps | 160 bytes | 58 | ~87.2 kbps | 100 kbps |
Look closely at G.729. Its raw payload is tiny-just 20 bytes per packet. But the overhead is 58 bytes. That means nearly 75% of the bandwidth consumed by a G.729 call is just protocol headers! This is why G.729 is excellent for saving WAN costs when bandwidth is scarce, but inefficient if you have plenty of pipe available. Using G.711 on a fast link is often better than G.729 on a slow link because the complexity and CPU load of compression aren’t worth the marginal bandwidth savings if you already have capacity.
The Hidden Costs: Encryption and IPv6
The calculations above assume a basic, unencrypted IPv4 environment. Real-world enterprise networks rarely look like that. Two major factors inflate true bandwidth further:
SRTP Encryption: Secure Real-time Transport Protocol (SRTP) encrypts the voice payload. While it doesn’t change the header size significantly, it can prevent certain optimizations like Compressed RTP (cRTP) from working effectively. Some implementations also add authentication tags that increase the payload size slightly.
IPv6 Headers: If your network runs IPv6, the IP header jumps from 20 bytes to 40 bytes. That’s an extra 20 bytes of overhead per packet. For a G.729 call, this increases the overhead-to-payload ratio even further. Always check if your endpoints and routers are dual-stack or pure IPv6, as this silently eats into your capacity.
VPNs: If your VoIP traffic traverses a site-to-site VPN, you’re adding another layer of encapsulation. IPSec, for example, can add 40-60 bytes of additional overhead per packet. Suddenly, that "light" G.729 call is consuming close to 50 kbps instead of 31 kbps.
How to Calculate Your Actual Needs
Don’t guess. Use a structured approach to determine the bandwidth required for your specific setup.
- Identify the Codec: Check what your phones and PBX negotiate. Internal calls usually default to G.711 or G.722. External calls might force G.729 or Opus.
- Determine Packetization: Confirm the ptime setting. 20 ms is standard, but some systems use 30 ms to save bandwidth.
- Count Concurrent Calls: Peak usage matters more than average. Plan for your busiest hour, not your quietest.
- Apply the Multiplier: Use the table above to find the one-way bandwidth. Multiply by two for full-duplex (bidirectional) traffic.
- Add Signaling Overhead: SIP signaling (INVITE, ACK, BYE messages) uses TCP or UDP and consumes bandwidth too. While small per message, bursts during registration storms or call setups can spike usage. Reserve about 10-15% extra for signaling and control traffic.
- Include QoS Buffer: Quality of Service queues need headroom to prioritize voice over data. Never run your link at 100% utilization.
For example, if you expect 20 concurrent G.711 calls:
- Per call bidirectional: ~174 kbps (87 x 2)
- Total media bandwidth: 20 × 174 kbps = 3,480 kbps (3.48 Mbps)
- Add 15% for signaling/QoS: 3.48 Mbps × 1.15 ≈ 4 Mbps
If you only planned for 64 kbps per call, you’d have calculated 2.56 Mbps. You would have been short by over 1.4 Mbps, leading to congestion.
Tips to Reduce Overhead Impact
If you are constrained by bandwidth, you can’t eliminate overhead entirely, but you can mitigate its impact.
Use Compressed RTP (cRTP): cRTP shrinks the 40-byte IP/UDP/RTP header down to 2-4 bytes by exploiting the fact that many header fields remain constant between packets. This is huge for low-bandwidth links. However, both ends must support it, and it breaks easily if there is packet loss or if NAT/firewalls interfere. Test thoroughly before deploying.
Increase Packetization Interval: Moving from 20 ms to 30 ms reduces packets per second from 50 to roughly 33. This cuts overhead bandwidth by ~33%. The trade-off is increased latency. For internal LAN calls, stick to 20 ms. For remote sites with tight bandwidth, try 30 ms.
Prioritize Codecs Wisely: Don’t force G.729 everywhere. If your WAN link is robust, let endpoints negotiate G.711 or Opus. The higher quality is worth the extra bandwidth, and you avoid the CPU cost of compression/decompression. Save G.729 for truly constrained connections.
Avoid Unnecessary Encryption on Trusted Networks: Within a secure, segmented VLAN, you might skip SRTP to save payload space and allow cRTP. Only enforce encryption on untrusted public internet legs.
Common Pitfalls in Bandwidth Planning
I’ve seen three mistakes repeat themselves in VoIP deployments:
- Ignoring Upstream Bottlenecks: Most businesses have asymmetric internet (high download, low upload). VoIP needs consistent upstream bandwidth. A 100 Mbps download plan with only 10 Mbps upload can only support ~50 G.711 calls, not hundreds.
- Forgetting Jitter Buffers: When bandwidth is tight, jitter buffers expand to smooth out packet arrival times. This adds latency. If you’re right at the limit of your bandwidth, quality degrades not just due to loss, but due to excessive buffering delay.
- Assuming Static Rates for Opus: Opus is adaptive. It might drop to 10 kbps during silence or spike to 120 kbps during complex audio. Planning for the average (40 kbps) is risky. Plan for the peak or use a conservative upper bound (e.g., 80 kbps) to ensure stability during dynamic content.
Frequently Asked Questions
Does using Wi-Fi affect VoIP bandwidth calculations?
Yes, significantly. Wi-Fi introduces Layer 2 overhead that differs from Ethernet. 802.11 frames include management and control overheads that vary with signal strength and interference. Additionally, Wi-Fi retransmissions due to packet loss consume bandwidth without delivering new data. Always reserve extra headroom (at least 20-30%) for wireless VoIP deployments compared to wired ones.
Is G.729 always more efficient than G.711?
Not necessarily. While G.729 uses less bandwidth (approx. 31 kbps vs. 87 kbps for G.711), it requires significant CPU resources for compression and decompression. On modern hardware with ample bandwidth, G.711 is often preferred for its lower latency, superior audio quality, and lack of transcoding artifacts. G.729 is best reserved for scenarios where bandwidth is strictly limited, such as satellite links or congested WAN circuits.
How does IPv6 impact VoIP bandwidth?
IPv6 headers are 40 bytes long, compared to 20 bytes for IPv4. This adds 20 bytes of overhead per packet. For a G.729 call, this increases the per-packet size from 78 bytes to 98 bytes, raising the bandwidth requirement from ~31 kbps to ~39 kbps. For G.711, it raises the requirement from ~87 kbps to ~95 kbps. Always account for this larger header size if your network is IPv6-only.
What is Compressed RTP (cRTP) and should I use it?
cRTP compresses the IP, UDP, and RTP headers from 40 bytes down to 2 or 4 bytes. This drastically reduces overhead, especially for small payloads like G.729. However, it requires support on both endpoints and intermediate devices. It can fail if packet loss occurs, forcing a fallback to uncompressed headers. Use cRTP on low-bandwidth WAN links where every bit counts, but test extensively to ensure compatibility with your firewalls and session border controllers.
How much bandwidth do I need for 100 concurrent Opus calls?
Opus is variable, but for planning purposes, assume a conservative average of 40 kbps per direction (80 kbps bidirectional). For 100 calls, that’s 8 Mbps of media bandwidth. Add 15-20% for signaling and safety margins, bringing the total to approximately 9.5-10 Mbps. If you use high-definition Opus profiles, budget closer to 12-15 Mbps to be safe.
Write a comment