← Back to eksmanager.io
Case Study

Five years, six clusters, and one accident

A financial institution's Kubernetes estate, run for five years across a datacentre, Azure, and now AWS. The clusters always ran. Keeping them upgraded depended on one person — this is what it took to change that.

The estate

A financial institution's Kubernetes estate: five clusters on Azure, a datacentre until two and a half years ago, and since this year a cluster on AWS. Five years managing it; two and a half years of the platform running it.

The clusters are the environments — dev, QA, production — so a change has to survive the same path on every one of them, and a difference between clusters is a difference between what was tested and what is live.

Every cluster had its own core stack — its own ArgoCD, its own ingress, its own certificate handling — assembled by whoever built it, at whatever versions were current that week. They drifted apart, because nothing held them together except me knowing which was which.

Elasticsearch is the clearest example. They wanted better logging, so I learned ECK, deployed it, and maintained it. Nobody else knew much about it. That was true of the stack generally: I was the only person who could upgrade it, or the clusters underneath it.

The accident

Then a serious accident put me out of action with no notice, and no clear date for coming back. Three months, possibly longer.

The clusters kept running. What I had built by then was semi-automated — enough that the day-to-day looked after itself. What stopped was the upgrades. Nobody else could upgrade the core stack, or the clusters underneath it.

In a regulated organisation that is not an inconvenience. It is patches not being applied, and it is a line on a risk register: the estate depends on one person's knowledge, and that knowledge is not written down anywhere a colleague could use at two in the morning.

What replaced it

So from 2024Q2 the knowledge went into a system instead.

One core stack, identical on every cluster — ArgoCD and Argo Workflows for delivery, Traefik and ExternalDNS for ingress and DNS, ECK for search and logging — all at versions tested together rather than assembled independently. Three upgrades a year, on a fixed cadence, with an audit trail of what changed and when, and the change itself raised and approved in Jira.

TLS is handled outside the clusters on purpose. Let's Encrypt certificates are issued per environment and stored in the vault or secrets manager, kept fresh by Terraform and EventBridge, and deployed into the cluster by the platform. Because they are versioned where they are stored rather than held only inside a cluster, rolling one back is a decision rather than an incident.

The datacentre went in the same period — migrated off first, then shut down. Both pieces of work pointed the same way: a simpler, cloud-managed estate, with fewer things behaving differently depending on where they happened to run.

Twelve hours, to under three

12 hrs
an upgrade by hand — and a weekend was not reliably long enough
< 30 min
the core stack, now
< 2 hrs
the cluster upgrade itself, unattended

An upgrade used to take twelve hours by hand, and a weekend was not reliably long enough — which is why one of them was postponed to the next window rather than finished in a hurry.

The core stack now takes under thirty minutes. The cluster upgrade itself runs under two hours, unattended, because Elasticsearch is scaled down onto its retained volumes before the nodes roll rather than having terabytes shuffled between them as they go.

What the record looks like

No production outage caused by the platform or by an upgrade, across six clusters and every release since 2024Q2.

The worst it has done is a gap in ECK logging. Worth stating plainly, because a platform that has never surprised anyone is either very new or being described by someone who wasn't watching.

Consistency, and access no wider than the job

The throughline was never the tooling. It was two things.

Clusters that are the same as each other, because the clusters are the environments — so a difference between them is a difference between what was tested and what is live. And access no wider than the job requires, because a platform that needs standing administrative rights to upgrade itself has moved the risk rather than removed it.

Both are easier to deliver on AWS. The permission model is more granular — scoped roles per function, credentials issued rather than held, and an account boundary that makes the blast radius of a mistake something you can draw on a diagram. That granularity is what turns least privilege from something you document into something you implement.

What it turned out to be worth

Not the setup. Anyone can stand a cluster up, and with today's tooling they can assemble a plausible stack in an afternoon.

It's the second year that's hard — knowing what changed between one release and the next, which combination has actually been run together, and what breaks when it hasn't. And it's being able to answer an auditor without reconstructing it from memory.

The accident was the point. A platform that only one person can upgrade isn't a platform; it's a dependency with a pulse.

Moving off a datacentre, or standardising an estate you inherited?

Get in touch