← Back to results

Site Reliability Engineer

Posted 7 days ago
San DiegoLast seen 5 days ago

Summary

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines. The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.