Setting maxUnavailable: 0 on a Deployment is the thing you do when you have decided that dropping requests during a deploy is not acceptable. It reads like a promise. Kubernetes will not take a pod away until a replacement is ready, so there is always a full complement of healthy pods, so nobody should ever get an error.
I wanted to know whether that holds, so I put four replicas behind a cloud load balancer, pointed twenty keep-alive clients at it, and triggered a rolling update while counting every failure. The answer is that it does not hold, which I half expected. What I did not expect was that the fix everybody recommends barely moved the number, and that the thing which actually fixed it was somewhere else entirely.
The setup
A small Python HTTP server, four replicas, a type: LoadBalancer Service on DigitalOcean Kubernetes 1.36.3, two nodes. The clients run HTTP/1.1 with keep-alive and reuse their connections, because that is what real clients do, and because a load generator that opens a fresh connection per request will hide most of this.
The deploy strategy is the cautious one:
strategy:
rollingUpdate: {maxSurge: 1, maxUnavailable: 0}
Each run sends about 19,000 requests over 110 seconds and triggers one rolling update 25 seconds in. The only thing that changes between runs is how the application handles being shut down.
Before any of it, the control. The same load, the same duration, no rollout:
requests=19380 ok=19380 failed=0
Zero failures across nineteen thousand requests, which means anything the rollout produces is the rollout's fault rather than background noise. I almost skipped this step and I should not have, because without it every number below is unreadable.
Three attempts, none of which worked
Attempt one, the naive application. It handles SIGTERM by exiting immediately, which is roughly what a process does if you never think about it:
requests=19302 ok=19251 failed=51
33 RemoteDisconnected
13 ConnectionRefusedError
4 ConnectionResetError
Fifty-one failures, and on a repeat run, thirty-five. Note what they are not. Not one of them is a 5xx. Every single failure is at the connection level, which is why a monitoring setup that counts HTTP status codes will tell you the deploy was clean.
Attempt two, proper graceful shutdown. The server stops accepting new connections on SIGTERM, lets in-flight requests finish, then exits a few seconds later. This is what every framework means when it advertises graceful shutdown.
requests=19392 ok=19367 failed=25
25 RemoteDisconnected
Half the failures gone, which is real progress, and twenty-five still there.
Attempt three, the fix everybody recommends. Add a preStop hook so the pod sits still for a moment before shutdown begins, giving the endpoint time to be withdrawn:
lifecycle:
preStop:
exec: {command: ["/bin/sleep", "5"]}
requests=19308 ok=19288 failed=20
Twenty-five became twenty. For a change that is supposed to be the answer, that is not an answer.
Where the failures actually were
This is the point where the experiment got useful, because a result that refuses to move means the explanation is wrong rather than the fix being insufficient.
I had a watcher recording, four times a second, which pod addresses were in the Service's EndpointSlice, and the pods were logging the moment they received SIGTERM. Lining those two up gives the gap between a pod being told to stop and that pod being taken out of rotation.
| median gap | |
|---|---|
no preStop
|
+0.57s |
preStop: sleep 5 |
-0.06s |
Positive means the endpoint was still live after SIGTERM had already arrived, which is the race everyone describes. It is real, it is about half a second wide, and preStop closes it completely. The endpoint now leaves rotation just before the process is signalled.
So preStop did its job perfectly and twenty requests still failed. Those twenty cannot be new traffic arriving at a dead pod, because no new traffic was being sent there.
Plotting the surviving failures against the SIGTERM timestamps answers it immediately.
Every failure lands 3.0 seconds after a SIGTERM. Three seconds is exactly how long my graceful handler waits before calling os._exit. The failures are not happening when the pod is told to stop, they are happening when the pod actually stops.
The load balancer is holding established keep-alive connections to that pod. Refusing new connections does nothing about them, and a preStop sleep does nothing about them either. They sit in a pool, perfectly healthy as far as the pool is concerned, until the process exits and severs them mid-use.
What actually fixed it
Once the problem is stated that way the fix is obvious, and it is not a Kubernetes setting at all. The application has to get itself out of the connection pool rather than waiting to be removed from it. On HTTP/1.1 that means telling the client to stop reusing the connection:
if draining.is_set():
self.send_header("Connection", "close")
self.close_connection = True
Keep serving normally, but close each connection after its response. The pool drains itself over the next few seconds. Then exit, with terminationGracePeriodSeconds set high enough that nothing kills you first.
requests=25884 ok=25884 failed=0
requests=25871 ok=25871 failed=0
Zero, twice, across more than fifty thousand requests. The full progression, same cluster, same load, same rollout:
| failures | |
|---|---|
| no rollout (control) | 0 |
| exits on SIGTERM | 51 |
| graceful shutdown | 25 |
| graceful + preStop | 20 |
| graceful + preStop + closes keep-alives | 0 |
What I got wrong
Two things, and the first is the one that nearly cost me the post.
I built this expecting preStop to be the ending. The planned article was the familiar one about the endpoint removal race, with a satisfying drop to zero when the sleep goes in. When attempt three came back at twenty instead of zero I spent a while assuming five seconds was not long enough and that I should try fifteen, which would have been a slow way to learn nothing. The timing data was already sitting in a file I had not looked at.
The second is that I nearly did not run the control. It felt redundant, since obviously a steady load against a steady deployment does not fail. Had the baseline come back at twenty failures, the entire comparison would have been noise and I would have published a story about a bug that was in my load generator. One extra run of the thing where nothing happens is the cheapest insurance available.
There is also a smaller one worth passing on. I deleted the cluster and watched the droplets disappear, and the load balancer stayed. A type: LoadBalancer Service provisions a real one, and it is a separate object with a separate bill that outlives the thing that asked for it. It did clear on its own after about three minutes in my case, but if you are tearing down a test cluster, go and look at the load balancer list afterwards rather than assuming.
What to take from it
If you run a Deployment behind any load balancer that reuses connections, which is all of them, then maxUnavailable: 0 and a graceful shutdown and a preStop hook together still leave you dropping requests on every deploy. Not many, and never as a 5xx, so you will not see them unless you are counting connection errors.
Three things have to be true. The endpoint has to leave rotation before the process is signalled, which is what preStop buys you. In-flight requests have to finish, which is what graceful shutdown buys you. And established idle connections have to be closed by you rather than severed by your exit, which is the one nobody mentions and the only one that took my number to zero.
The check is cheap. Point a keep-alive client at your service, roll a deploy, and count failures rather than status codes. If the number is not zero, the connections are the first place to look.














