Engineering Lead · Ollion

Amit Sharma — Cloud & Platform Engineering Leader

Zero-downtime cloud migrations, SRE-grade Kubernetes platforms, and the engineering programmes that ship them.

I design resilient cloud platforms and lead the engineering programs that ship them. From SRE-grade Kubernetes to zero-downtime migrations, I own delivery end-to-end across GCP, AWS, and Azure.

What I Stand For

Principles that shape how I work

01

Zero downtime is a constraint, not a goal

Teams that achieve zero downtime design for it from day one. It's not something you bolt on at the end — it shapes every architectural decision, from dual-run strategies to rollback triggers.

02

Migrations are org problems first

The technical challenges are rarely what derails a migration program. Misaligned stakeholders, unclear ownership, and InfoSec teams looped in too late — these are the real risks.

03

Platform engineering is about enabling autonomy

A good platform team makes itself invisible by enabling other teams to move fast safely. The goal isn't to run Kubernetes — it's to make the 100 teams using it not need you.

Selected Work

Case studies with real outcomes

Latest Thinking

Writing from two angles

Insights

Technical & industry perspective

All →
Platform Engineering 5 min

Debugging Across Layers Is the Platform Skill Nobody Hires For

Job descriptions list Kubernetes, Linux, networking, storage and virtualization as separate skills. The fault you get paged for lives in the seams between them. Here's how I work a problem that crosses layers, and why I think that's the skill to hire for.

Migration Strategy 4 min

Map What Exists Before You Draw What Should

Every migration and modernisation programme I've run started with a target architecture someone had already drawn. The ones that worked threw it away after discovery. Here's why the dependency map is the real architecture artefact, and how I build one.

Migration Strategy 4 min

Migration and Improvement Can Share a Window, Not a Change

Every migration customer eventually says it: if we're taking downtime anyway, let's fix a few things while we're in there. I say no every time. You can share the window and the error budget. You cannot share the change.

Platform Engineering 6 min

You Don't Own a Kubernetes Platform Until You Own the Layer Beneath It

Running self-managed Talos Kubernetes on Nutanix, with no cloud provider under it, changed how I design platforms on managed clouds too. Here's what the layer beneath teaches you.

Platform Engineering 5 min

Idle Capacity Is the Most Expensive Thing on Your Platform

I built my first capacity tooling in 2015, bursting HPC jobs to the cloud when the on-prem queue backed up. A decade later, on much bigger fleets, the lesson is unchanged: idle capacity and queue wait are the same number seen from two sides, and both should be on the dashboard next to latency.

AI in DevOps 5 min

Give the Agent the Keyboard, Keep the Approval

I've rolled out AI tooling across a delivery team and I'm building an agentic workflow that drafts Terraform and opens pull requests. The rule that's held up: agents may author, humans must approve, and the review bar doesn't move.

Perspectives

AI, society & the human dimension

All →

From the Lab

Self-hosting and systems experiments

All →

More on how I got here in about, or see the full career history.

Certifications

AWS & Google Cloud certified

Solutions Architect – Professional

Amazon Web Services

Professional Cloud Architect

Google Cloud

Professional Cloud Security Engineer

Google Cloud

Generative AI Leader

Google Cloud