← Back to results

etcd jobs in San Diego

$149,800 – $262,200 · Posted 5 days ago

Senior Staff Software Engineer leading design and delivery of multi-tenant, distributed data platform services on Kubernetes at scale. You will mentor mid-level and junior engineers, architect highly available stateful workloads (databases, message queues, caches), and integrate AI into infrastructure and operational workflows.… The role requires 8+ years building production distributed systems, deep expertise in Kubernetes and at least one hyperscaler (AWS, Azure, GCP), and hands-on experience operating services like Postgres, Kafka, or Redis at scale with explicit HA, failover, and durability targets.

San DiegoLast seen 4 days ago
$92,400 – $148,800 · Posted 1 month ago

A Senior Site Reliability Engineer at SHEIN will own and operate mission-critical, large-scale distributed systems (Kubernetes, Kafka, Elasticsearch, Redis, APISIX, Nginx) running 24/7/365, participating in on-call rotations and driving fast incident response using AI-assisted log analysis and anomaly detection. The role requires strong software engineering expertise in Python or Go, deep Linux and networking knowledge, and hands-on experience with observability platforms (Prometheus, Grafana) and configuration management tools.… Responsibilities include designing resilient monitoring and alerting infrastructure, automating operational workflows to eliminate toil, capacity planning, and collaborating with global teams to improve system reliability and performance.

San DiegoLast seen 1 day ago
$108,000 – $180,000 · Posted 1 month ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper).… Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 9 days ago