Sitemap

We Cut Our Kubernetes Pods by 60% and Doubled Traffic Capacity

6 min readJan 30, 2026

--

The “Everything is Fine” Moment

You know that feeling when your monitoring dashboards are green, but something still feels… off? That was us three months ago.

One of our critical services was running with 26 pods. Everything was “stable.” But here’s the thing — during normal business hours, most of those pods were just sitting there, barely breaking a sweat. Then, during traffic spikes, we’d see weird hiccups. Response times would jump. Some requests would timeout. It was like having a sports car that struggled to start.

My manager asked the obvious question: “Why are we paying for 26 pods if most of them are doing nothing?”

I didn’t have a good answer.

The Uncomfortable Truth

Press enter or click to view image in full size

This was our “stable” system — struggling at 5K requests per second with 26 pods running

Looking at this graph now, I can see what we couldn’t see then. We were over-provisioned and under-optimized at the same time. It’s like wearing a winter coat in summer while still feeling cold — something was fundamentally wrong with our approach.

The “What If” Conversation

Over coffee one morning, our team started asking dangerous questions:

“What if we just… reduced the minimum pods to 10?”

“What if we let Kubernetes actually do its job?”

“What if we’re solving the wrong problem?”

These are the kinds of conversations that either lead to brilliant solutions or career-limiting moves. Spoiler alert: it worked out.

What Actually Changed (And Why It Mattered)

Discovery #1: We Were Choking Our Own JVM

Our JVM was configured with MaxRAMPercentage=75. Sounds reasonable, right? Give the JVM most of the container's memory, leave a little for the OS.

Wrong.

Here’s what we didn’t account for: metaspace, direct buffers, thread stacks, GC overhead, and all the other stuff that lives outside the heap. We were basically asking our JVM to fit into a box that was too small, then wondering why it kept crashing into the walls.

The fix? Drop it to 65%.

-XX:MaxRAMPercentage=65

Suddenly, no more random OOMKills. It’s like when you stop wearing shoes that are one size too small — you don’t realize how much pain you were in until it’s gone.

Discovery #2: Our Database Had 1,300 Connections (Yes, Really)

This one hurt. We had 26 pods, each with a Hikari connection pool of 50 connections.

26 × 50 = 1,300 database connections.

For context, our database was screaming. We thought we needed more database resources. Turns out, we just needed to stop attacking it with a thousand connections.

Old settings:

maximumPoolSize: 50
minimumIdle: 25

New settings:

maximumPoolSize: 20
minimumIdle: 10

Now, even at peak load (30 pods), we’re using 600 connections instead of 1,300. Our DBA literally sent us a thank you message.

Discovery #3: Kubernetes Can Scale Fast (If You Let It)

Here’s where things got interesting. We completely rewrote our HPA strategy around one simple idea: when traffic comes, react immediately. When traffic drops, wait and see.

minReplicas: 10
maxReplicas: 30
scaleUp:
stabilizationWindowSeconds: 0 # React NOW
policies:
- type: Pods
value: 4 # Add 4 pods per minute
- type: Percent
value: 100 # OR double the pods

scaleDown:
stabilizationWindowSeconds: 300 # Wait 5 minutes
policies:
- type: Percent
value: 25 # Remove max 25% per minute

Why the difference? Because users don’t care if we run a few extra pods for 5 minutes. They absolutely care if our service is slow because we didn’t scale up fast enough.

The Load Test That Made Us Believers

We ran the test. Held our breath. Watched Grafana like it was the Super Bowl.

Press enter or click to view image in full size

Same service, different configuration — 10.5K requests per second, smooth as butter

The numbers speak for themselves:

  • Before: 5,000 req/s (struggling)
  • After: 10,500 req/s (comfortable)
  • Response time: Consistent 897ms at 90th percentile
  • Errors: Basically none
  • Baseline pods: 10 (down from 26)

We literally doubled our capacity while using 60% fewer resources during normal operation.

I had to run the test three times because I didn’t believe it.

What This Actually Looks Like in Production

A month after deployment, here’s what changed:

Monday morning (normal traffic):

  • Running with 10–12 pods
  • Handling ~3,000 req/s
  • Response times around 200–300ms
  • Cost: ~$X/day (actual numbers redacted, but significantly less)

Friday afternoon (peak traffic):

  • Scales to 28–30 pods in 2 minutes
  • Handling 8,000+ req/s
  • Response times still around 500–800ms
  • Cost: Only pay for what we use

Late night (minimal traffic):

  • Back down to 10 pods
  • Handling ~500 req/s
  • System stays warm and ready

No OOM kills. No CPU throttling. No panicked Slack messages.

The Lessons (That I Wish Someone Had Told Me)

1. “Stable” doesn’t mean “optimized”

Our old setup was stable like a boulder is stable. Sure, it wasn’t crashing, but it wasn’t doing much of anything useful either.

2. Your JVM needs breathing room

That extra 10% headroom (from 75% to 65% MaxRAMPercentage) saved us countless debugging sessions. Give your runtime space to exist.

3. More connections ≠ better performance

We thought our database was the bottleneck. Turns out we were the bottleneck, flooding it with unnecessary connections. Sometimes the best optimization is doing less.

4. Aggressive scale-up is a feature, not a bug

Yes, you might over-provision for a few minutes. But that’s infinitely better than under-provisioning during a traffic spike. Users remember slow; they don’t remember “slightly more expensive for 5 minutes.”

5. Test your assumptions

I was convinced we needed 26 pods. I was wrong. Don’t be married to your initial decisions.

If You’re Thinking About Doing This

Here’s my honest advice: Start small. Don’t change everything at once.

Week 1: Fix your JVM settings. Test them. Make sure nothing explodes.

Week 2: Adjust your connection pools. Monitor database performance.

Week 3: Update your HPA configuration. Watch it like a hawk.

Week 4: Run load tests. Lots of them.

Week 5: Deploy to production during a low-traffic period.

And for the love of all that is holy, have a rollback plan.

The Unexpected Benefits

Beyond the obvious cost savings and performance improvements, something else happened: we stopped worrying.

Before, every traffic spike was a potential incident. We’d watch the dashboards nervously, wondering if today was the day we’d need to manually scale.

Now? The system just handles it. Kubernetes does its job. We can focus on building features instead of babysitting infrastructure.

That peace of mind is worth more than any cost savings.

Your Turn

If you’re running services on Kubernetes and something feels off — even if your dashboards are green — trust that instinct. Maybe you’re over-provisioned. Maybe your connection pools are too large. Maybe your JVM is suffocating.

Start measuring. Start testing. Start asking “what if?”

The worst case? You learn something and go back to your original config.

The best case? You double your capacity and cut your costs.

Want to dive deeper into the technical details? I’ve documented our exact HPA configuration, JVM flags, and load testing methodology in our internal wiki. If there’s enough interest, I might write a follow-up on the monitoring and alerting setup we built around this.

Hit me up in the comments if you’ve done something similar, or if you’re thinking about it and have questions. I’ve made enough mistakes in this process that I can probably save you from a few of them.

P.S. — Our DBA is still sending us thank you notes for the connection pool optimization. Best relationship improvement we’ve had all year.

--

--