Reliability engineering for teams running on AWS, GCP and Kubernetes. I design production systems, cut cloud spend, and build the observability and automation that let engineers ship without fear.
Architect, build and operate production-grade Kubernetes on AWS & GCP — from cluster design to self-service tooling for your engineers.
Cut cloud bills without cutting reliability. Usage visibility, right-sizing and architectural changes — the kind that took ~70% off a real bill.
Prometheus, Thanos, Grafana and CloudWatch platforms with meaningful SLIs/SLOs — so you find problems before your users do.
Terraform, AWS CDK, GitLab CI and GitHub Actions. Repeatable provisioning and delivery pipelines your team can actually own.
High-availability design, autoscaling GPU/AI workloads, Kafka data pipelines and migrations that improve DR and scale.
Harden infrastructure for standards like ISO 27001 — reducing risk while keeping delivery moving.
I have read the Privacy Policy.