← Back to results

chaos-engineering jobs in San Diego

Posted 7 days ago

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines.… The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.

San DiegoLast seen 5 days ago
Posted 23 days ago

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production services running on AWS and Kubernetes. Responsibilities include designing highly available infrastructure with Terraform and managed AWS services, building CI/CD pipelines with GitHub Actions and Argo CD, defining SLOs and implementing observability with Prometheus and Grafana, leading incident response, and automating operational work with Python, Go, or Bash.… The role requires 3–8 years of hands-on SRE or DevOps experience, production Kubernetes expertise, strong AWS and infrastructure-as-code knowledge, and proficiency with observability and deployment strategies.

San DiegoLast seen 21 days ago
Posted 1 month ago

This Staff Platform Engineer role owns operational strategy and technical direction for enterprise GitLab (3,000+ users) and JFrog Artifactory platforms, leading cloud-native transformation and Kubernetes migrations. The position requires 7+ years of enterprise-scale platform operations, 5+ years each with self-managed GitLab and Artifactory in high-availability deployments, and 3+ years of deep Kubernetes expertise.… Responsibilities include designing multi-region active-active systems, establishing cloud-native best practices, mentoring team members, and influencing organizational standards for platform engineering and developer experience. Required expertise spans Linux kernel/networking, PostgreSQL high availability, AWS/Azure/GCP, Kubernetes operators, service mesh (Istio), infrastructure-as-code (Terraform), observability (Prometheus), and production automation in Python, Go, and Bash.

San DiegoLast seen 1 month ago