I Was Wrong About Kubernetes Swap, And It Actually Matters Now
Adil Sher
Author
I spent the last three years telling anyone who'd listen that swap in Kubernetes was a terrible idea. "Just buy more RAM," I'd say confidently, usually while configuring another overprovisioned node pool. Then I got a project involving AI agent workloads, and suddenly my certainty evaporated. It turns out the landscape has genuinely shifted, and I needed to update my mental model.
The breaking point came when we deployed our first batch of sandboxed AI agents on our GKE cluster. These things spin up a massive memory footprint just to initialize, hundreds of megabytes sitting there waiting for the next prompt, then go almost completely idle. We were paying for that idle memory across dozens of pods, and our node utilization numbers looked embarrassing. That's when the latest Kubernetes blog post on swap-backed node density landed in my feed, and I realized the old constraints I was thinking about no longer applied.
Why I Used to Hate Swap in Kubernetes
My skepticism wasn't irrational. There were genuine, historical reasons Kubernetes didn't support swap well. Under cgroup v1, the kernel couldn't properly track swap separately from physical memory, so a container's actual memory usage became unpredictable. A pod claiming 256MB could silently page hundreds of megabytes to disk, breaking isolation and making resource planning impossible. On top of that, paging to spinning disk was genuinely terrible for latency, you'd see tail latencies spike into the seconds range.
So the advice was sensible at the time: avoid swap, provision enough physical RAM, and accept the density ceiling. For traditional workloads, this worked fine. You sized your nodes, set your memory requests, and moved on.
What Changed: cgroup v2 and NVMe
The technical foundation shifted in two ways. First, cgroup v2 gives us separate accounting for swap versus physical memory. A container can be limited to, say, 256MB of physical RAM and 512MB of swap, and the kernel actually respects those boundaries independently. The accounting is clean and predictable again.
Second, Local NVMe SSDs are cheap now. When swap backs to fast NVMe instead of spinning disk, the latency penalty becomes negligible for workloads that don't constantly fault. You're not replacing your active working set, you're giving the kernel a pressure valve for burst memory and idle state.
The Actual Density Gains Are Substantial
The Kubernetes team benchmarked this across three real workload categories, and the numbers are worth taking seriously. CI/CD kernel builds cut their memory requirements in half with no latency increase. Headless browser sandboxes (the exact use case we're running) achieved 2-3x density improvements depending on the runtime isolation level.
Here's what caught my attention: the Python sandbox results showed 200% density improvement. That means a node that could run 80 isolated Python runtimes without swap now runs 240. That's not marginal, that's the difference between needing three nodes and needing one.
My Take: This Is Contextual But Valuable
I'm not saying swap is universally correct now. It's not. For latency-critical services with sustained high throughput, you still want your working set in physical RAM. If your workload is memory-bound (constantly hitting swap), you've sized your cluster wrong.
But for bursty, agentic, and batch workloads with idle-heavy characteristics? This changes the economics. We're at an inflection point where AI workloads are driving cluster composition, and they don't behave like traditional services. They spike, they plateau, they go quiet. Swap absorbs that pattern elegantly.
One thing I'd want to validate before rolling this into production is tail latency. The benchmarks show geometric mean latency is fine, but I'd want to see p99 and p999 numbers. A 40% increase in execution time when the active set is forced into swap (as noted in the kernel build test) is acceptable, but only if it's predictable and not the common case.
The Implementation Question
The practical question for my team is whether to enable this cluster-wide or selectively. I'm leaning toward selective adoption, enable swap on specific node pools designated for AI workloads and batch jobs, keep it off for services where latency predictability is non-negotiable.
What I want to know from you: are you running production Kubernetes workloads with swap enabled yet? What's your experience with the latency profile? I'm specifically curious about real-world p99 numbers and whether the density gains hold up when you're actually running heterogeneous workloads together.
Source: This post was inspired by "Scaling Kubernetes Workloads with Node Swap" by Kubernetes Blog. Read the original article