Experience supporting startups, enterprise platforms, and globally distributed engineering teams.
9 years building and modernizing production infrastructure, reduce operational noise, and build reliable platforms through Kubernetes, automation, observability, and cloud engineering.
- Contract SRE / Platform work
- Infrastructure modernization
- Kubernetes migrations
- Reliability consulting
- Modernize legacy infrastructure
- Reduce alert fatigue
- Migrate to Kubernetes
- Improve observability and incident response
- Reduce cloud spend without sacrificing reliability
Palo Alto Networks • Hopin • DeepDyve • Policy Networks • Startups & SMBs
Platform Kubernetes • Helm • Terraform
Observability Prometheus • Grafana
Cloud & Security GCP • Cloudflare
Data PostgreSQL • Redis • Elasticsearch
Automation GitHub Actions • Python • Go • Bash
- Led observability modernization from legacy based Nagios to Prometheus/Grafana across production environments
- Reduced alert fatigue through smarter routing, automation, and operational tuning
- Migrated legacy VM-based infrastructure into containerized Kubernetes platforms
- Designed resilient multi-cluster environments for disaster recovery and operational continuity
- Reduced cloud and infrastructure spend through modernization and platform optimization
- 🚀 Nagios → Prometheus Migration (case study underway)
- Kubernetes Platform Modernization (case study underway)
- 📉 Alert Fatigue Reduction Initiative (case study underway)
- 🔐 Disaster Recovery & Multi-Cluster Resilience (case study underway)
🔗 LinkedIn: Zett Stai


