Site Reliability Engineer
ExpiredSan DiegoLast seen 21 days ago
Summary
The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production services running on AWS and Kubernetes. Responsibilities include designing highly available infrastructure with Terraform and managed AWS services, building CI/CD pipelines with GitHub Actions and Argo CD, defining SLOs and implementing observability with Prometheus and Grafana, leading incident response, and automating operational work with Python, Go, or Bash. The role requires 3–8 years of hands-on SRE or DevOps experience, production Kubernetes expertise, strong AWS and infrastructure-as-code knowledge, and proficiency with observability and deployment strategies.
DevOps / InfrastructureAWSKubernetesPythonTerraformApache KafkaArgo CDBashBlue Green DeploymentCanary DeploymentCapacity PlanningChaos EngineeringCI CDDisaster RecoveryDynamodbEksGitHub ActionsGitopsGoGrafanaHelmIncident ResponseInfrastructure As CodeObservabilityOpentelemetryPrometheusRdsRolling ReleaseS3Service Level ObjectivesService MeshesSlos