we are looking for a Senior Site Reliability Engineer (SRE).
As a Senior Site Reliability Engineer on the R&D Infrastructure team in our Tel Aviv Office, youll play a vital role in building, scaling and maintaining high-scale infrastructure across on-premise cloud, public cloud and rapidly growing AI/ML Kubernetes environments. You will push Linux to its limits, writing software and automation to eliminate manual tasks and solving performance bottlenecks across the stack.
As a Senior Site Reliability Engineer on the R&D Infrastructure team in our Tel Aviv Office, youll play a vital role in building, scaling and maintaining high-scale infrastructure across on-premise cloud, public cloud and rapidly growing AI/ML Kubernetes environments. You will push Linux to its limits, writing software and automation to eliminate manual tasks and solving performance bottlenecks across the stack.
Requirements:
7+ years of experience managing, scaling and troubleshooting large-scale distributed Linux environments in production.
Deep understanding of Linux system internals and network protocols (TCP/IP, DNS, HTTP, gRPC), together with hands-on experience of edge/CDN services such as Fastly, Cloudflare, Akamai or CloudFront.
Hands-on experience with Infrastructure as Code (IaC) and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD or Jenkins.
Production experience managing containerized environments using Kubernetes and Docker.
Solid programming skills in at least one modern language (Go, Python or Rust).
Bonus points if you have:
Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK/logging).
Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem.
7+ years of experience managing, scaling and troubleshooting large-scale distributed Linux environments in production.
Deep understanding of Linux system internals and network protocols (TCP/IP, DNS, HTTP, gRPC), together with hands-on experience of edge/CDN services such as Fastly, Cloudflare, Akamai or CloudFront.
Hands-on experience with Infrastructure as Code (IaC) and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD or Jenkins.
Production experience managing containerized environments using Kubernetes and Docker.
Solid programming skills in at least one modern language (Go, Python or Rust).
Bonus points if you have:
Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK/logging).
Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem.
This position is open to all candidates.








