← Back to results

site-reliability-engineering jobs in San Diego

$155,300 – $258,800 · Posted 2 days ago

Lead enterprise Azure cloud platform strategy and operations for a financial services firm, overseeing compute, storage, networking, containers, and disaster recovery capabilities. Drive adoption of cloud-native architectures, managed services, and Azure AI/OpenAI to improve engineering productivity. Manage cloud governance, security compliance, FinOps practices, and cost optimization across the organization. Serve as senior escalation point and strategic partner with Microsoft, using Jira and ServiceNow to coordinate delivery and service management.

San DiegoLast seen today
$177,300 – $265,900 · Posted 10 days ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues. Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 8 days ago
$108,000 – $180,000 · Posted 14 days ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper). Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 12 days ago