Most network load balancing failures come from four places: a wrong algorithm, broken health checks, firewall rules blocking probes, or session persistence set incorrectly. Verify the listener and algorithm first, then confirm every backend passes health checks from the balancer itself, then check firewall and DNS. Fix in that order.
Symptoms
- One backend takes 80% of requests while others sit idle or barely warm.
- Users get logged out mid-session, or their shopping cart empties between clicks.
- Intermittent 502, 503, or 504 responses under normal load.
- Health check dashboard flaps servers between healthy and unhealthy every few seconds.
- New connections stall or time out even though backend servers respond directly.
Common Causes
Wrong or mismatched balancing algorithm
Round Robin on long-lived connections concentrates load. Least Connections on very short requests can behave like Random. The algorithm has to match the actual traffic pattern.
Health checks that lie
Probing TCP/80 tells you the socket is open, not that the app works. A shallow check keeps a broken node in rotation and starves the healthy ones.
Session persistence misconfigured
Sticky sessions off when the app needs state causes logouts. Sticky on when it doesn't pins users to one node and defeats balancing entirely.
Firewall or security group blocking probes
The balancer's health-check source IPs or the VIP itself get blocked by a host firewall, WAF, or upstream ACL. Backends look dead from the balancer, fine from a jump host.
DNS, routing, or stale ARP
Clients resolve to an old VIP after a failover. ARP caches point at the previous active node. Traffic goes to a black hole for minutes.
Step-by-Step Fix
- Confirm the balancer is actually listening and forwarding
On Windows NLB, run `nlb.exe query` on the cluster host. On HAProxy or Nginx, check `sudo systemctl status haproxy` (or nginx) and confirm the process is running. Verify the listener with `ss -tlnp` or `netstat -an | grep LISTEN` and match it against the ports your clients use. If the socket isn't bound, nothing downstream matters. - Review the balancing algorithm against real traffic
Round Robin assumes requests cost roughly the same. Least Connections works better for variable request times. IP Hash (or Source affinity) suits apps holding server-side state without a shared session store. Pick the one that matches your workload, then reload the config. Don't leave the default because it shipped that way. - Ensure all backends are alive and not saturated:
Ping will not tell you everything. Test app port: curl -I http://backend:8080/healthz on the load balancer host, not your laptop. Linux backends: CPU pressure with top/htop; memory with free -m; connection count with ss -s. - Don't test a port, test the app:
Make the health check call a real endpoint like /health or /status that actually hits the database or cache that the app relies on. Adjust the interval, timeout, and unhealthy threshold so that a slow GC pause won't take out a node, but an app outage will within 15-30 seconds. - Make sure the firewall paths are open for probes and traffic:
Allow the balancer source IPs (or security group) to each backend on the health-check port and service port. Make sure all host firewalls are open too: iptables -L -n (Linux); Get-NetFirewallRule (Windows). A backend that responds from the same subnet but not from the balancer subnet is nearly always a firewall or routing issue. - Match session persistence with the app:
For local session state, turn on stickiness by cookie (preferably) or source IP. For a Redis or database session store, turn stickiness off to transparently fail over if needed. Test this by tailing the access logs on all backends while a single user clicks through your app; requests should either pin to one node or disperse, as you have chosen, not dance between them. - Verify DNS TTL, VIP ownership, ARP:
For DNS balancing, use a low TTL (30-60s) to accelerate failover. When sharing a VIP, verify one node has it: ip addr show on Linux, Get-NetIPAddress on Windows. After a failover, clear upstream ARP, or inject a gratuitous ARP to clear the ARP table in switches. Outdated ARP tables are a common cause of post-failover black holes. - Reload firmware and config, and observe:
Vendor firmware/package updates during a scheduled window – older releases of HAProxy, F5 TMOS, or NLB have known bugs that cause these issues. After your changes, open the balancer and backend access logs side by side for 15 min with load; monitor request counts per backend, not total request count.
Common load balancing symptoms mapped to likely cause and first fix
| Symptom | Likely cause | First action |
|---|---|---|
| One backend gets most traffic | Round Robin with long-lived connections, or IP Hash on few source IPs | Switch to Least Connections and reload |
| Users logged out randomly | Session persistence off, app stores state locally | Enable cookie-based stickiness |
| Backends flap healthy/unhealthy | Probe timeout too tight or shallow TCP check | Increase timeout, use HTTP check on /health |
| 502 or 504 under load | Backend saturated or keepalive mismatch | Check backend CPU and align keepalive timeouts |
| Traffic dies after failover | Stale ARP or DNS cache | Send gratuitous ARP; lower DNS TTL |
| Balancer sees backend down, direct curl works | Firewall blocks probe source IP | Allow balancer subnet on backend firewall |
Prevention
- Alert on per-backend request rate variance, not just total 5xx count.
- Keep health-check endpoints in version control and test them in CI.
- Run a scheduled failover drill each quarter and time the recovery.
- Document which apps need sticky sessions and why, next to the balancer config.
FAQ
Should I use Round Robin or Least Connections?
Round Robin is great for stateless, short, uniform requests such as serving static assets. Least Connections is the better choice for mixed workload types, as it takes into consideration the number of active connections on each backend. If one or two backends are slower than the others, Least Connections will prefer them until they fill up.
My healthchecks are all green, but users are receiving errors?
It sounds like the probe is likely too shallow. A TCP check just verifies the socket is open, and an HTTP check of / could be returning a cached page. You should point your probe at an actual endpoint that can cause the app to hit your database, cache, or any subsequent API it depends on.
Do I require sticky sessions on a contemporary app?
Session state should be stored on a local server only. Sessions stored in Redis, Memcached, or a database also don't require stickiness; they are safer without it. Stickiness by cookie is safer than source IP, because corporate NATs and cellular carriers hide hundreds of users behind one address and badly skew distribution.