← Back to results

prometheus jobs in San Diego

$92,400 – $148,800 · Posted 4 days ago

A Senior Site Reliability Engineer at SHEIN will own and operate mission-critical, large-scale distributed systems (Kubernetes, Kafka, Elasticsearch, Redis, APISIX, Nginx) running 24/7/365, participating in on-call rotations and driving fast incident response using AI-assisted log analysis and anomaly detection. The role requires strong software engineering expertise in Python or Go, deep Linux and networking knowledge, and hands-on experience with observability platforms (Prometheus, Grafana) and configuration management tools. Responsibilities include designing resilient monitoring and alerting infrastructure, automating operational workflows to eliminate toil, capacity planning, and collaborating with global teams to improve system reliability and performance.

San DiegoLast seen 2 days ago
$150,000 – $220,000 · Posted 7 days ago

Shield AI seeks a Staff Engineer, DevOps to own and maintain the build system for autonomous aircraft (XBAT/VBAT), with deep expertise in C++ and CMake. You will improve developer velocity by reducing build/test cycle times, integrating CI pipelines with HIL and simulation environments, and implementing monitoring and dashboards for build system performance. The role requires 7+ years of industry experience (or 5+ with advanced degree), proficiency with containers/orchestration, scripting, CI/CD tools, and cloud operations across Azure/AWS/GCP. You'll collaborate with software engineers to resolve failures, set best practices, and leverage agentic AI solutions to identify root causes and accelerate development cycles.

San DiegoLast seen 4 days ago
Posted 9 days ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code. Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 7 days ago
Posted 9 days ago

Build and operate the data and ML infrastructure powering an AI platform for materials science, owning both sides: data pipelines that ingest and curate large-scale scientific output into training-ready formats, and model packaging, serving, monitoring, and CI/CD systems that move models safely from research to production across customer environments. You will design data ingestion and transformation workflows, implement validation and quality gates, package and version models with reproducible builds, run models through batch and online inference with safe rollout and rollback, monitor for drift and degradation, and build observability and internal tooling for engineering and science teams. The role requires 6+ years shipping production software with deep expertise in data systems, ML infrastructure, containers, orchestration, and observability.

San DiegoLast seen 7 days ago
$125,000 – $160,000 · Posted 11 days ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response. Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 9 days ago
$125,000 – $160,000 · Posted 12 days ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes. Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 10 days ago
$108,000 – $180,000 · Posted 12 days ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper). Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 10 days ago
Posted 12 days ago

This Staff Platform Engineer role owns operational strategy and technical direction for enterprise GitLab (3,000+ users) and JFrog Artifactory platforms, leading cloud-native transformation and Kubernetes migrations. The position requires 7+ years of enterprise-scale platform operations, 5+ years each with self-managed GitLab and Artifactory in high-availability deployments, and 3+ years of deep Kubernetes expertise. Responsibilities include designing multi-region active-active systems, establishing cloud-native best practices, mentoring team members, and influencing organizational standards for platform engineering and developer experience. Required expertise spans Linux kernel/networking, PostgreSQL high availability, AWS/Azure/GCP, Kubernetes operators, service mesh (Istio), infrastructure-as-code (Terraform), observability (Prometheus), and production automation in Python, Go, and Bash.

San DiegoLast seen 10 days ago