The Great Container Migration: How We Survived Moving 200 Microservices from Swarm to Kubernetes

When Docker Swarm Started Feeling Like a Comfortable Prison

Three years ago, our Docker Swarm cluster was humming along beautifully. We had 47 services running across 12 nodes, deployments took seconds, and our monitoring dashboard glowed a reassuring green most nights. Then we hit that magical point where “simple” becomes “simplistic,” and what once felt elegant started feeling like we were trying to run a Formula 1 race in a golf cart.

The first warning sign wasn’t dramatic. Our data pipeline team casually mentioned they needed more granular resource controls for their ML workloads. Then the security team started asking uncomfortable questions about network policies. Finally, our newest engineer looked at our deployment scripts and asked, “Why can’t we just use Helm charts like everyone else?” That’s when I knew we were living in the past.

Docker Swarm had worked well during our scrappy startup days. The learning curve was gentle, the mental model straightforward, and it rarely woke us up at night. But as we scaled past 100 services and added compliance requirements, we found ourselves implementing increasingly baroque workarounds for problems that Kubernetes solved out of the box. Sometimes the tool that got you here isn’t the tool that gets you there.

Planning the Great Escape Without Setting Everything on Fire

The temptation was to rip the band-aid off quickly. Spin up a shiny new EKS cluster, migrate everything over a weekend, and emerge victorious on Monday morning. This is exactly the kind of thinking that leads to resume-generating events and heated Slack conversations with your CTO at 2 AM.

Instead, we took the boring approach that actually works. We spent two months building a comprehensive migration plan, starting with our least critical services as guinea pigs. Our staging environment became a Kubernetes playground where we could break things safely and learn from our mistakes when the stakes were low. We documented every gotcha, every configuration difference, and every “wait, how did we handle that in Swarm?” moment.

The real breakthrough came when we realized we didn’t need to migrate everything at once. We set up ingress controllers that could route traffic between our Swarm services and new Kubernetes deployments, effectively running a hybrid setup for months. This gave us the luxury of moving services one at a time, validating each migration thoroughly before moving to the next. It felt slow at the time, but it saved us from the chaos that comes with trying to debug 47 broken services simultaneously.

The Devil in the Configuration Details

If you’ve never migrated from Docker Swarm to Kubernetes, here’s what they don’t tell you in the blog posts: volume mounts will make you question your life choices. Swarm’s simple volume syntax transforms into a maze of persistent volumes, storage classes, and claims that require you to understand the difference between ReadWriteOnce and ReadWriteMany in ways you never wanted to.

Our logging service migration became a week-long odyssey because we discovered that our Swarm setup had been quietly mounting the host’s Docker socket into containers. Kubernetes took one look at that configuration and basically said “absolutely not.” We ended up redesigning our entire logging architecture around proper sidecar containers and centralized collection, which was ultimately better but required rewriting deployment configs for 23 different services.

Environment variable injection was another delightful surprise. Swarm let us get away with sloppy practices around secret management that Kubernetes simply wouldn’t tolerate. We had to properly implement ConfigMaps and Secrets, which forced us to finally solve the “how do we manage configuration across environments” problem we’d been postponing for months. The migration revealed technical debt we didn’t even know we had accumulated.

When the Rubber Met the Road

The moment of truth came during our first major service migration. Our payment processing API had been running flawlessly on Swarm for two years, handling thousands of transactions daily with rock-solid reliability. Moving it to Kubernetes felt like performing surgery on a perfectly healthy patient while they were awake and asking why you were doing this to them.

We deployed the Kubernetes version alongside the Swarm instance, using feature flags to gradually shift traffic. For three weeks, we ran both versions in parallel, comparing metrics, response times, and error rates down to the millisecond. The Kubernetes deployment actually performed slightly better once we tuned the resource requests and limits properly, but the real win was the operational visibility we gained through proper health checks and readiness probes.

The final cutover happened on a Tuesday at 2 PM, when traffic was predictably moderate. We flipped the switch, held our breath, and watched our monitoring dashboards like hawks. Everything worked. No alerts fired. Transaction processing continued without a blip. It was almost anticlimactic, which is exactly what you want when migrating critical infrastructure.

Life After the Migration

Six months later, our Kubernetes cluster is managing 187 services across multiple namespaces with the kind of operational sophistication that would have been impossible in our Swarm days. Resource utilization is more efficient, scaling is more predictable, and our deployment pipelines finally support the advanced patterns our developers had been requesting for years.

But the real victory isn’t technical. We can now hire experienced engineers without having to explain why we’re still using orchestration tools from 2016. Our interview process no longer includes the awkward conversation about whether candidates are willing to learn our “unique” deployment approach. We’re using industry-standard tooling with industry-standard practices, which turns out to matter more than I initially thought.

The migration taught us that sometimes the biggest risk is staying put. Swarm worked fine for what it was, but it was limiting our ability to evolve. Kubernetes brought complexity, yes, but it also brought capabilities we didn’t even know we needed until we had them. Sometimes you have to make things harder in the short term to make them easier in the long term.

I’d be curious to hear from other teams who’ve made similar migrations. What surprised you most about the process? What would you do differently if you had to do it again? Drop me a line or share your own war stories in the comments.

You may also like