← Back to results

site-reliability-engineering jobs in San Diego

$142,000 – $213,000 · Posted 6 days ago

Staff Platform Engineer responsible for designing, building, and operating large-scale cloud and hybrid infrastructure platforms supporting mission-critical systems. Will develop infrastructure-as-code frameworks, establish standardized deployment patterns, build CI/CD pipelines and developer tooling, and collaborate with engineering and security teams to embed security controls and improve system reliability.… Requires 8+ years in platform engineering, DevOps, or SRE roles with deep hands-on experience in cloud infrastructure, networking, identity, and storage systems, plus strong expertise in IaC tools like Terraform and Ansible, scripting/programming languages, and Linux environments.

San DiegoLast seen 4 days ago
Posted 14 days ago

As a Software Engineer on the Platform Engineering team, you will architect and deliver cloud-native developer platforms and services that power ResMed's digital health ecosystem. You'll own the full lifecycle from design through production operations, building observability and reliability into services while participating in 24/7 on-call incident management.… The role requires strong fundamentals in at least one programming language (Java, Python, Go, TypeScript, JavaScript, React), hands-on experience with AWS and cloud-native technologies (Kubernetes, Docker, serverless), and a developer-first mindset focused on improving tooling and automation for internal engineering teams.

San DiegoLast seen 12 days ago
$143,000 – $215,000 · Posted 14 days ago

Senior Platform Engineer to design, build, and operate a cloud-native compute platform on AWS and Kubernetes. You will own platform security, scalability, and evolution through strategic initiatives and migration efforts, lead incident response and troubleshooting, and provide technical leadership and mentorship.… Requires 8+ years in platform/SRE/infrastructure engineering with deep AWS expertise, production Kubernetes experience, hands-on CI/CD (GitHub, GitHub Actions, Argo CD), Infrastructure as Code (Terraform/CDK), Kubernetes networking knowledge, and production observability/monitoring skills.

San DiegoLast seen 12 days ago
$230,000 – $290,000 · Posted 20 days ago

Lead a Reliability Platform Engineering team at Affirm responsible for building observability, risk management, and operational intelligence systems that help engineers manage production reliability at scale. You'll translate operational challenges into technical requirements, develop scalable reliability capabilities leveraging AI and automation, and drive alignment across Platform Engineering, SRE, Infrastructure, and product teams.… The role requires 7+ years of backend/full-stack engineering experience with 2+ years of engineering leadership, deep expertise with observability tools and SRE practices, and strong programming skills in Python, Kotlin, Java, or similar languages.

San DiegoLast seen 19 days ago
$114,400 – $114,400 · Posted 21 days ago

Senior Systems Engineer responsible for administration and operations of a global cloud SaaS infrastructure serving millions of users. The role requires expert-level Linux and systems administration skills, with hands-on responsibility for availability, performance monitoring, and troubleshooting across web, database, and network layers in production environments.… Must have SRE/DevOps experience managing large-scale server infrastructure, proficiency with configuration management tools (Ansible, Puppet, Chef), monitoring platforms (Nagios, Splunk), and scripting languages (Python, Perl, or JavaScript). Knowledge of databases (MySQL, PostgreSQL, Oracle), load balancing (F5, NGINX), and networking technologies is essential.

San DiegoLast seen 19 days ago
$123,000 – $185,000 · Posted 26 days ago

Staff Platform Engineer at Shield AI will design, build, and operate highly available infrastructure platforms supporting mission-critical systems across cloud and on-prem environments. Responsibilities include improving CI/CD pipelines, developer platform tooling, and automation frameworks; implementing scalable solutions for compute, networking, identity, and storage; and partnering with security teams to embed controls and compliance into platform architectures.… The role requires 8+ years of hands-on platform engineering, DevOps, or SRE experience with deep expertise in Infrastructure as Code (Terraform, Ansible), Linux systems, distributed systems, and scripting/programming languages (Python, Go, Bash, PowerShell).

San DiegoLast seen 24 days ago
Posted 27 days ago

Staff Platform Engineer at Shield AI responsible for designing, building, and operating highly available cloud and hybrid infrastructure platforms supporting mission-critical systems. The role spans infrastructure-as-code development, CI/CD pipeline enhancement, developer tooling, and platform automation to accelerate engineering productivity.… Requires 8+ years of platform engineering, DevOps, or SRE experience with deep hands-on expertise in cloud infrastructure, networking, identity, storage, and infrastructure automation across Linux environments. Expected to partner with security teams on compliance and DevSecOps practices, establish scalable platform patterns, and lead complex platform initiatives across large-scale engineering organizations.

San DiegoLast seen 21 days ago
$205,489 – $244,000 · Posted 1 month ago

Lead a team to improve engineering efficiency and product quality in software development for 5G and RAN optimization. Design and implement scalable, production-grade software solutions using distributed systems and AI principles.… Serve as a technical expert and staff-level contributor, mentoring development teams, driving architecture decisions, and translating telecom requirements into robust code while collaborating with site reliability engineers and cross-functional teams.

San DiegoLast seen 1 month ago
$120,001 – $160,000 · Posted 1 month ago

Design, develop, and deploy a hybrid container/virtualization platform serving enterprise workloads across on-prem and cloud environments. Integrate core platform services (identity, PKI, DNS, logging, monitoring, secrets management) and support GPU-enabled AI/ML workloads.… Establish CI/CD pipelines, Infrastructure-as-Code automation, and standard operating procedures for configuration management, verification, and patch management. Lead cross-functional teams on containerization, DevSecOps, incident response, and ATO compliance efforts.

San DiegoLast seen 1 month ago
$69,300 – $158,000 · Posted 1 month ago

As an AWS Cloud Security Engineer, you'll conduct hands-on security assessments and engineering for enterprise AWS environments, reviewing Organizations, accounts, landing zones, IAM, networking, and security services to identify gaps and strengthen cloud posture. You'll work with AWS-native services (IAM, CloudTrail, GuardDuty, Security Hub, KMS, etc.), develop Terraform Infrastructure as Code, establish secure configuration baselines, create security policies and guardrails, and translate technical requirements into scalable, repeatable solutions.… You'll collaborate with cloud architects, security engineers, platform teams, and client leadership to deliver implementation-ready documentation and cloud governance improvements. The role requires 3+ years of AWS cloud engineering, cloud security, or DevSecOps experience, with hands-on expertise in multi-account AWS environments, VPC networking, security monitoring services, and Infrastructure as Code.

San DiegoLast seen 1 month ago
$155,300 – $258,800 · Posted 1 month ago

Lead enterprise Azure cloud platform strategy and operations for a financial services firm, overseeing compute, storage, networking, containers, and disaster recovery capabilities. Drive adoption of cloud-native architectures, managed services, and Azure AI/OpenAI to improve engineering productivity.… Manage cloud governance, security compliance, FinOps practices, and cost optimization across the organization. Serve as senior escalation point and strategic partner with Microsoft, using Jira and ServiceNow to coordinate delivery and service management.

San DiegoLast seen 1 month ago
$177,300 – $265,900 · Posted 1 month ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues.… Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 11 days ago
$108,000 – $180,000 · Posted 1 month ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper).… Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 15 days ago