TL;DR — Key Takeaways

  • Modern platforms automate provisioning, deployment and scaling, yet many production incidents still depend on manual investigation, bridge calls and tribal knowledge.
  • Kubernetes orchestrates workloads but does not by itself guarantee application-level resilience or business availability.
  • The next stage of platform engineering is automating repeatable recovery decisions, testing failure regularly and measuring application recovery rather than infrastructure recovery alone.

I find that I cannot stop asking myself this question… How did we automate almost everything in software delivery, yet when production fails, we still throw humans at the problem?

Think about what modern platform engineering has accomplished. We provision infrastructure with code. CI/CD? We build pipelines that can deploy hundreds of times a day. We use GitOps to keep environments consistent. Workloads? Ours can automatically scale based on demand. Kubernetes? We can spin up an entire Kubernetes cluster in minutes. That is incredible progress.

Then a production database fails at 2:13 a.m.

The Slack channel lights up. Someone starts a bridge call. Someone else starts digging through a runbook that hasn’t been updated in a year. And everyone waits for the one engineer who “knows how this thing works.”

How is that still our operating model?

We’ve spent the last decade teaching our platforms how to deploy software. We haven’t spent nearly as much time teaching them how to respond when production breaks. That’s the next maturity curve for platform engineering – not more deployment automation, but intelligent operational automation.

One of the biggest misconceptions I see? Organizations that equate modern infrastructure with resilient infrastructure. They’re not the same thing, not at all. Kubernetes is fantastic at orchestrating workloads. Better than any platform before, it can restart a pod, replace a failed node, and reconcile desired state.

But Kubernetes doesn’t know whether your customers can still place an order, process a payment, or complete a transaction. That’s not a knock on Kubernetes. It’s simply not its job. The problem is we’ve started treating orchestration as if it were a complete reliability strategy. It isn’t.

Orchestration answers the question, “Where should this workload run?”

Reliability answers a different question: “How do I keep this workload available when something inevitably fails?”

These are not the same problem. In modern environments, availability shouldn’t be tied to a server, a cluster, or even a cloud. It should follow the workload.

Here’s the question every platform team should ask. When I speak with engineering teams, I tell them to forget buzzwords for a minute. Forget “cloud native.” Forget “self-healing.” Forget “high availability.” Instead, ask one simple question, “If something breaks in production, what should happen next – and how much of it still depends on people?”

If the honest answer is, “Well…someone gets paged, we jump on a call, and then we figure it out,” your platform isn’t as automated as you think it is. It’s not a criticism. It’s an opportunity.

The goal isn’t to remove engineers from the process. It’s to eliminate the routine operational decisions they’ve already made hundreds of times before. Engineers should spend their time solving new problems – not repeating the same recovery steps every time a server, node, or workload fails.

So…What Should You Do Differently on Monday?

Here’s where I’d start.

Find your 2 a.m. processes.

Make a list of everything your team does manually during a production incident. Not deployments. Not upgrades. Failures. If only one person knows how to recover a critical workload, that’s not expertise – it’s technical debt. Recovery shouldn’t live inside someone’s head. It should be built into the platform.

Measure application recovery – not infrastructure recovery.

Most teams know exactly how long it takes to deploy a release. Far fewer know how long it takes for a critical application to recover from a real failure. Your users don’t care how quickly a pod restarted. They care how quickly they can get back to work. Measure what they experience. Better yet, start measuring how many recovery decisions still require human intervention. That may be the best indicator of platform maturity you have.

Test recovery as often as you test deployments.

Most organizations exercise their CI/CD pipelines constantly. In a controlled way, how often do you deliberately break production to see what happens? Is the answer, “Almost never?” – then I think you have found your next engineering project.

Stop assuming Kubernetes solved everything.

Kubernetes is one of the best orchestration platforms ever built. It’s also just one layer of your platform. Stateful applications, databases, storage, networking, and application dependencies all have their own failure modes. Make sure your operational strategy accounts for them. Modern platforms aren’t resilient simply because they run on Kubernetes. They’re resilient because every layer – from infrastructure to the application itself – is designed to recover intelligently when something breaks.

Eliminate decisions – not just steps.

The goal isn’t to make your outage runbook shorter. The goal is to make it unnecessary. The best platforms don’t just automate tasks – they automate decisions that have already been defined through policy and testing. That’s how you reduce downtime and reduce stress on your engineers at the same time.Here’s a simple test. If your most experienced engineer took two weeks off tomorrow, would your recovery process work exactly the same way? If the answer is “probably,” your platform still depends on tribal knowledge. Mature platforms don’t just automate tasks – they automate operational decisions.

The Next Frontier Isn’t Deployment. It’s Intelligent Operations

For years, we’ve measured platform engineering by one question: “How fast can we deploy?” I think there’s a better question for the decade ahead: “How intelligently can we recover?”

The answer has less to do with where an application runs than with how easily it can keep running when conditions change. Modern enterprises aren’t choosing one environment over another. They’re operating across on-premises infrastructure, virtual machines, Kubernetes clusters, public clouds, and edge environments – all at the same time. Reliability can no longer be tied to a particular server, cluster, or cloud. It has to follow the workload.

The best platform teams won’t be the ones that deploy the fastest. They’ll be the ones whose applications remain available no matter where those workloads are running – or where they need to move next.

That’s the next evolution of platform engineering: Building platforms where operational intelligence, not infrastructure, determines availability.

Frequently Asked Questions

Why isn’t Kubernetes enough for resilience?

Kubernetes can restart pods, replace nodes and reconcile desired state, but it does not inherently know whether customers can still complete a transaction or whether the wider application is functioning correctly.

What should platform teams measure instead of just deployment speed?

They should measure application recovery time, the number of recovery decisions requiring human intervention and whether users regain service quickly after a failure.

What does intelligent operations mean?

It means codifying known recovery decisions into policy and automation so routine incidents can be detected, diagnosed and remediated without waiting for an expert to intervene manually.

Share.
Leave A Reply