WAN Link Sizing for Multi-Site VoIP: A Practical Guide for Branch Offices

WAN Link Sizing for Multi-Site VoIP: A Practical Guide for Branch Offices

You bought the shiny new IP phones. You configured the PBX. But when the Monday morning rush hits, your branch office calls sound like they’re coming through a tin can. Why? Because you sized your WAN link based on marketing brochures, not physics. Most IT managers make the fatal mistake of assuming a 64 kbps G.711 call only uses 64 kbps of bandwidth. It doesn’t. Once you add headers, jitter buffers, and safety margins, that same call might eat up 90 kbps or more. If you have fifty people talking at once, your "plenty fast" circuit is suddenly choking to death.

The Hidden Cost of Every Packet

Let’s get real about what happens when you speak into a phone. The voice data isn't just raw audio; it's wrapped in layers of digital packaging. Think of it like shipping a single grape (the voice payload) in a massive crate (the packet). You pay for the crate, the label, and the truck space, not just the grape.

Codecs like G.711 and G.729 determine the size of the grape. G.711 is high quality but heavy (64 kbps payload). G.729 is compressed and light (8 kbps payload). But here’s the kicker: regardless of the codec, every packet needs an IP header, a UDP header, and an RTP header. On top of that, your Layer 2 framing (Ethernet or PPP) adds more weight.

If you use standard Ethernet with IPv4, you are adding roughly 40-50 bytes of overhead per packet. At a typical 20-millisecond packetization interval, that’s 50 packets per second. For a G.729 call, this overhead can triple the effective bandwidth usage compared to the raw codec bitrate. Ignoring this leads to under-provisioned links that fail exactly when you need them most.

Per-Call Bandwidth Estimates (Full Headers, No VAD)
Codec Payload (kbps) Overhead (kbps) Total Per Call (kbps)
G.711 64 ~26 ~90
G.729 8 ~26 ~34
iLBC 13.3 ~26 ~40

Calculating Your Busy Hour Concurrency

Knowing how much one call costs is half the battle. The other half is knowing how many calls happen at once. You don’t size for total users; you size for Busy Hour Call Attempts (BHCA).

A common rule of thumb from Cisco design guides suggests a 5:1 ratio of users to active calls during peak times. So, if you have 50 employees in a branch, expect around 10 simultaneous calls. But this is risky if your branch is a sales team making constant outbound dials. A better approach is using the formula:

Concurrent Calls = (BHCA × Average Call Duration in Seconds) / 3600

Let’s say your branch handles 600 calls in the busiest hour, and the average call lasts 180 seconds (3 minutes). That’s (600 * 180) / 3600 = 30 concurrent calls. See the difference? The 5:1 rule guessed 10. The math says 30. If you sized for 10, your network will melt down at noon.

The 33% Rule and Safety Margins

Once you have your total voice bandwidth requirement, do not buy a circuit that exactly matches that number. Voice traffic is bursty. It needs room to breathe. Industry best practices, reinforced by tools like Monocalc and Famoustec, recommend that voice traffic should never exceed 33% of your total WAN link capacity.

Why 33%? Because your WAN link also carries email, cloud backups, video conferencing, and random user browsing. If voice takes up 80% of the pipe, any small spike in data traffic will cause congestion. And when congestion hits, your QoS policies have to work overtime. If they fail even slightly, your voice quality drops instantly.

So, take your calculated voice bandwidth (e.g., 1.2 Mbps for 30 G.729 calls), add a 10-20% buffer for signaling and retransmissions, and then divide by 0.33. In our example, 1.2 Mbps / 0.33 ≈ 3.6 Mbps. You’d round up to a standard 5 Mbps circuit. This ensures that even if someone starts downloading a large file, your calls remain clear.

Cartoon robots packing tiny voice grape into huge header crate

Quality of Service (QoS): The Traffic Cop

Bandwidth is useless without control. You need Quality of Service (QoS) to prioritize voice packets over bulk data. Without QoS, a large email attachment sent during a call will compete directly with voice packets, causing jitter and delay.

The gold standard for enterprise branches is Low Latency Queuing (LLQ). LLQ creates a strict priority queue for voice. Voice packets jump the line immediately. Other traffic gets served fairly but only after voice is handled.

  • Voice Class: Assign DSCP EF (Expedited Forwarding). This gets strict priority.
  • Signaling Class: Assign DSCP AF31. SIP/H.323 messages need reliability but less urgency than media.
  • Data Class: Everything else. Use CBWFQ to share remaining bandwidth fairly.

Configure shaping on the egress interface of your branch router. If your link is 5 Mbps, shape to 5 Mbps. If you don’t shape, the router might try to push more data onto the link than it can physically handle, leading to tail-drop packet loss before QoS can even intervene.

SD-WAN and Encryption Overheads

If you’re using SD-WAN instead of traditional MPLS, add another layer of complexity. SD-WAN overlays often use encryption (like IPSec) and encapsulation (like GRE or VXLAN). These add significant header overhead-sometimes 20-40 bytes per packet more than standard IP.

This means your "per-call bandwidth" calculation must increase. A G.729 call that cost 34 kbps on a clean MPLS link might cost 45 kbps on an encrypted SD-WAN tunnel. Always test your actual throughput with tools like iPerf or built-in SD-WAN analytics before finalizing your QoS policy. Don’t trust the vendor’s generic calculator; trust your own measurements.

Cable-material traffic cop directing priority voice packets in cartoon

Common Pitfalls to Avoid

Even experienced engineers trip over these issues:

  1. Ignoring Signaling: SIP registration and keep-alives consume bandwidth. While small per device, hundreds of devices add up. Reserve 5-10 kbps per gateway for signaling.
  2. Forgetting Jitter Buffers: Routers and phones use buffers to smooth out arrival time variations. These buffers add latency. Keep one-way latency under 150ms. If your link is too full, queues build up, and latency spikes.
  3. Misclassifying Traffic: If your firewall or switch strips DSCP markings, your router won’t know which packets are voice. Ensure end-to-end QoS marking from the phone to the HQ.
  4. Static Sizing: Business changes. A new sales team doubles call volume. Review your BHCA metrics quarterly. What worked last year might be undersized today.

Practical Example: Sizing a 20-Person Branch

Let’s walk through a real scenario. You have a branch with 20 users. They use G.729 codecs. Historical logs show 400 BHCA with an average duration of 120 seconds.

  1. Calculate Concurrency: (400 * 120) / 3600 = 13.3 concurrent calls. Round up to 14.
  2. Calculate Raw Voice BW: 14 calls * 34 kbps/call = 476 kbps.
  3. Add Buffer: Add 20% for safety and signaling. 476 * 1.2 = 571 kbps.
  4. Apply 33% Rule: 571 kbps / 0.33 = ~1.73 Mbps minimum link speed.
  5. Select Circuit: A 2 Mbps circuit is tight. Choose a 5 Mbps circuit to allow for growth and data traffic.
  6. Configure QoS: Set LLQ to 600 kbps. Shape egress to 5 Mbps.

This methodical approach prevents the guesswork that plagues most deployments. It turns bandwidth planning from an art into a science.

Does Voice Activity Detection (VAD) really save bandwidth?

Yes, significantly. VAD (also known as silence suppression) stops sending packets when no one is speaking. Since humans are silent about 50% of the time during a conversation, VAD can cut your required bandwidth nearly in half. However, it requires compatible endpoints and careful configuration to avoid clipping words at the start of sentences.

What is the maximum acceptable latency for VoIP?

The general industry standard is one-way latency below 150 milliseconds. Above 200ms, conversations become awkward because of noticeable delays between speakers. Jitter should stay below 30ms, and packet loss should be under 1% for acceptable quality.

Can I use the same bandwidth calculation for video calls?

No. Video consumes vastly more bandwidth than voice. A standard HD video call can require 1-2 Mbps per stream. If your branch supports video conferencing, you must calculate video bandwidth separately and assign it its own QoS class, usually prioritized lower than voice but higher than bulk data.

How does SRTP affect bandwidth calculations?

Secure RTP (SRTP) adds authentication and encryption overhead. Typically, this adds about 10-20 bytes per packet. For high-volume sites, this can increase per-call bandwidth by 10-15%. Always account for this if security compliance mandates encrypted media streams.

What happens if I underestimate my WAN link size?

You will experience choppy audio, dropped calls, and long delays. More critically, if the link saturates, non-priority data traffic will starve, slowing down business applications. In severe cases, the router may drop voice packets entirely if the LLQ queue fills up due to excessive jitter or bursts.

VoIP bandwidth WAN link sizing branch office QoS codec overhead concurrent calls
Dawn Phillips
Dawn Phillips
I’m a technical writer and analyst focused on IP telephony and unified communications. I translate complex VoIP topics into clear, practical guides for ops teams and growing businesses. I test gear and configs in my home lab and share playbooks that actually work. My goal is to demystify reliability and security without the jargon.

Write a comment