Responsibilities:
Own operational responsibility for the application and platform layers.
Oversee and drive production operations end-to-end: deployments, maintenance, scaling, high availability and enhancements across AWS and GCP.
Operate Kubernetes in productions at scale across AWS EKS and GCP GKE
Manage all infrastructure as code.
Participate in on-call, lead and participate in incident management: triage, mitigation, and clear stakeholder communication during live incidents.
Own deployment tooling, troubleshooting, and performance tuning across compute, networking, and data layers.
Contribute to cost visibility and FinOps practices across cloud providers.
Own operational responsibility for the application and platform layers.
Oversee and drive production operations end-to-end: deployments, maintenance, scaling, high availability and enhancements across AWS and GCP.
Operate Kubernetes in productions at scale across AWS EKS and GCP GKE
Manage all infrastructure as code.
Participate in on-call, lead and participate in incident management: triage, mitigation, and clear stakeholder communication during live incidents.
Own deployment tooling, troubleshooting, and performance tuning across compute, networking, and data layers.
Contribute to cost visibility and FinOps practices across cloud providers.
Requirements:
Requirements
5+ years running production systems, hands-on.
Strong AWS experience (GCP a plus).
Production Kubernetes – networking, autoscaling, resource management, etc
Terraform (Terragrunt/Pulumi) as a daily tool, including modules and state.
Helm , GitOps with ArgoCD
GitHub and GitHub Actions; pipeline-as-code.
Solid Linux and networking: DNS, TLS, load balancing, VPC/routing.
Prometheus/Cortex, Grafana, Loki, or equivalents
Scripting in Python, Go, or Bash; comfortable reading application code.
Clear written English;
Experience in a distributed, global team.
Requirements
5+ years running production systems, hands-on.
Strong AWS experience (GCP a plus).
Production Kubernetes – networking, autoscaling, resource management, etc
Terraform (Terragrunt/Pulumi) as a daily tool, including modules and state.
Helm , GitOps with ArgoCD
GitHub and GitHub Actions; pipeline-as-code.
Solid Linux and networking: DNS, TLS, load balancing, VPC/routing.
Prometheus/Cortex, Grafana, Loki, or equivalents
Scripting in Python, Go, or Bash; comfortable reading application code.
Clear written English;
Experience in a distributed, global team.
Advantage
Development background.
Kafka, MQTT, or similar messaging at scale.
Data platform experience (Databricks, Spark,).
FinOps: cost allocation and unit economics.
Cloud security and compliance (SOC 2, CIS, CSPM).
Effective use of AI-assisted engineering tooling.
This position is open to all candidates.









