← Back to results

sli jobs in San Diego

Posted 7 days ago

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines.… The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.

San DiegoLast seen 5 days ago
Posted 15 days ago

Shield AI seeks an experienced SRE Lead to establish and mature reliability practices across Hivemind's cloud infrastructure and platform services. This hands-on technical role involves defining SLIs/SLOs, building monitoring and alerting systems, leading incident response, and driving root-cause analysis.… The SRE Lead will mentor teams, develop operational tooling in Python or Go, and partner with product and cloud engineering teams to embed reliability-first practices into system design and infrastructure provisioning.

San DiegoLast seen 13 days ago
Posted 16 days ago

Site Reliability Engineer responsible for monitoring system health, detecting issues before they escalate, and owning incident response and debugging end-to-end. You'll build logging and observability tooling, automate deployment pipelines, manage capacity planning, and drive platform reliability at scale.… The role requires 5+ years in SRE/DevOps, hands-on production incident response, and strong communication during live troubleshooting. You'll work with AWS, Terraform, Datadog, Python/Bash scripting, and modern infrastructure-as-code practices.

San DiegoLast seen 14 days ago
$150,100 – $225,100 · Posted 19 days ago

Software Engineer II responsible for operating and improving large-scale SQL, NoSQL, streaming, and caching platforms (Cassandra, Aurora, Aerospike, Kafka, Redis, DynamoDB, ElastiCache) with a focus on reliability, automation, and observability. Build self-service developer experiences using Terraform and infrastructure-as-code, participate in on-call incident response, and develop APIs and automated tests for infrastructure tooling.… Requires 3+ years in software engineering, database reliability engineering, SRE, or platform engineering, with hands-on Go experience, Terraform proficiency, AWS/GCP knowledge, and Kubernetes familiarity.

San DiegoLast seen 17 days ago
$150,100 – $225,100 · Posted 20 days ago

Software Engineer II responsible for operating and improving large-scale data platforms (SQL, NoSQL, streaming, caching) with a focus on reliability, automation, and self-service infrastructure. The role involves building observability tooling, participating in on-call incident response, troubleshooting database and infrastructure issues, and collaborating across teams to deliver reliable stateful services.… Required skills include Go, Terraform, API design, Kubernetes, and hands-on experience with databases like Cassandra, Aurora, Kafka, Redis, and DynamoDB on AWS or GCP.

San DiegoLast seen 18 days ago
$150,100 – $225,100 · Posted 20 days ago

Software Engineer II responsible for operating and improving large-scale data platforms (SQL, NoSQL, streaming, caching) with a focus on reliability, automation, and developer experience. The role involves building self-service infrastructure tooling, contributing to observability and SLO definitions, participating in on-call rotations, and automating infrastructure provisioning using Terraform.… Requires 3+ years of software or infrastructure engineering experience, production Go development, hands-on Terraform expertise, and deep knowledge of distributed systems concepts and operational practices.

San DiegoLast seen 19 days ago
$141,000 – $212,000 · Posted 23 days ago

Design, build, and operate scalable platform infrastructure across Azure, AWS, and private cloud environments using infrastructure-as-code, Kubernetes, and automation tooling. Develop reusable platform capabilities, CI/CD pipelines, and self-service infrastructure for engineering teams.… Own platform initiatives end-to-end from technical design through production operations, including deployment automation, configuration management, observability, and lifecycle management. Requires 7+ years of platform engineering or DevOps experience, strong hands-on proficiency with Terraform, Ansible, Python/Go, Kubernetes, and Linux systems administration.

San DiegoLast seen 21 days ago
$190,000 – $280,000 · Posted 29 days ago

Shield AI seeks an experienced Site Reliability Engineer to establish and mature the SRE function across Hivemind's cloud infrastructure and platform services. This hands-on leadership role involves defining reliability targets (SLIs/SLOs), improving observability and incident response practices, investigating complex failures, and driving adoption of reliability-first engineering practices.… The SRE Lead will mentor teams, develop operational automation tooling, and manage the multi-quarter SRE roadmap while working closely with Cloud Engineering and product teams to enhance system resilience and operational excellence.

San DiegoLast seen 27 days ago
$177,300 – $265,900 · Posted 1 month ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues.… Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 5 days ago
Posted 1 month ago

This Staff Platform Engineer role owns operational strategy and technical direction for enterprise GitLab (3,000+ users) and JFrog Artifactory platforms, leading cloud-native transformation and Kubernetes migrations. The position requires 7+ years of enterprise-scale platform operations, 5+ years each with self-managed GitLab and Artifactory in high-availability deployments, and 3+ years of deep Kubernetes expertise.… Responsibilities include designing multi-region active-active systems, establishing cloud-native best practices, mentoring team members, and influencing organizational standards for platform engineering and developer experience. Required expertise spans Linux kernel/networking, PostgreSQL high availability, AWS/Azure/GCP, Kubernetes operators, service mesh (Istio), infrastructure-as-code (Terraform), observability (Prometheus), and production automation in Python, Go, and Bash.

San DiegoLast seen 1 month ago