← Back to results

alerting jobs in San Diego

$195,000 – $255,000 · Posted 2 days ago

Senior Software Engineer responsible for designing and operating Affirm's analytical data platform, including secure data access patterns, CI/CD integration, self-service automation, and observability infrastructure. The role requires hands-on expertise with cloud data platforms (Snowflake, Databricks, BigQuery), containerization, API integrations, and Python/SQL backend work.… You'll participate in on-call coverage, triage production incidents, and help build next-generation capabilities including semantic layers and AI-native data tooling while setting high engineering standards across the platform.

San DiegoLast seen today
$79,040 – $120,640 · Posted 4 days ago

Design, build, and operate secure, scalable DevOps platforms at enterprise scale, with deep expertise in Kubernetes, Jenkins, Docker, and CI/CD infrastructure. Manage containerized workloads, implement monitoring and observability, provision and operate AWS infrastructure, and ensure platform reliability through incident response and capacity planning.… The role requires 3+ years of IT experience (or 5+ without a degree) with strong Linux administration, hands-on Kubernetes and Docker expertise, Jenkins infrastructure administration at scale, and CI/CD platform knowledge. This is a hands-on engineering role focused on platform reliability, security hardening, and developer enablement.

San DiegoLast seen 2 days ago
$190,000 – $280,000 · Posted 11 days ago

Senior Staff Lead Site Reliability Engineer to establish and mature SRE practices across Hivemind's cloud infrastructure and platform services. You will define reliability targets (SLIs/SLOs), build observability systems, lead incident response and root-cause analysis, mentor teams on reliability-first practices, and develop automation to reduce manual operational work.… The role requires 7+ years in SRE or infrastructure engineering, hands-on experience operating production services in AWS or equivalent cloud environments, expertise with containerized/distributed systems, infrastructure-as-code, and operational tooling development in Python or Go.

San DiegoLast seen 9 days ago
Posted 19 days ago

Shield AI seeks an experienced SRE Lead to establish and mature reliability practices across Hivemind's cloud infrastructure and platform services. This hands-on technical role involves defining SLIs/SLOs, building monitoring and alerting systems, leading incident response, and driving root-cause analysis.… The SRE Lead will mentor teams, develop operational tooling in Python or Go, and partner with product and cloud engineering teams to embed reliability-first practices into system design and infrastructure provisioning.

San DiegoLast seen 17 days ago
$155,380 – $221,416 · Posted 21 days ago

Director-level role managing infrastructure and operations for a healthcare systems company. Responsible for leading on-premises data center and enterprise facilities infrastructure, overseeing Windows/Linux servers, virtualization (VMware), SAN/NAS storage, disaster recovery, cloud platforms (Azure, AWS, GCP), networking (LAN/WAN, SD-WAN, firewalls, DNS, DHCP), identity and access management, and infrastructure monitoring.… Requires 10 years of subject-matter expertise, 8 years supervisory experience, and deep knowledge of IT service management, infrastructure automation, and compliance.

San DiegoLast seen 19 days ago
$150,100 – $225,100 · Posted 23 days ago

Software Engineer II responsible for operating and improving large-scale SQL, NoSQL, streaming, and caching platforms (Cassandra, Aurora, Aerospike, Kafka, Redis, DynamoDB, ElastiCache) with a focus on reliability, automation, and observability. Build self-service developer experiences using Terraform and infrastructure-as-code, participate in on-call incident response, and develop APIs and automated tests for infrastructure tooling.… Requires 3+ years in software engineering, database reliability engineering, SRE, or platform engineering, with hands-on Go experience, Terraform proficiency, AWS/GCP knowledge, and Kubernetes familiarity.

San DiegoLast seen 21 days ago
$150,100 – $225,100 · Posted 24 days ago

Software Engineer II responsible for operating and improving large-scale data platforms (SQL, NoSQL, streaming, caching) with a focus on reliability, automation, and self-service infrastructure. The role involves building observability tooling, participating in on-call incident response, troubleshooting database and infrastructure issues, and collaborating across teams to deliver reliable stateful services.… Required skills include Go, Terraform, API design, Kubernetes, and hands-on experience with databases like Cassandra, Aurora, Kafka, Redis, and DynamoDB on AWS or GCP.

San DiegoLast seen 22 days ago
$150,100 – $225,100 · Posted 24 days ago

Software Engineer II responsible for operating and improving large-scale data platforms (SQL, NoSQL, streaming, caching) with a focus on reliability, automation, and developer experience. The role involves building self-service infrastructure tooling, contributing to observability and SLO definitions, participating in on-call rotations, and automating infrastructure provisioning using Terraform.… Requires 3+ years of software or infrastructure engineering experience, production Go development, hands-on Terraform expertise, and deep knowledge of distributed systems concepts and operational practices.

San DiegoLast seen 23 days ago
$105,780 – $189,347 · Posted 26 days ago

The MLOps Engineer II designs, develops, and operates scalable machine learning infrastructure and deployment pipelines on AWS, working with data scientists and cloud engineers to productionize ML models. Responsibilities include building ML pipelines using SageMaker, Lambda, Step Functions, and S3; implementing CI/CD pipelines with GitHub and AWS tools; developing Infrastructure-as-Code using CloudFormation, Terraform, or AWS CDK; and ensuring production ML systems are reliable, secure, and cost-efficient.… The role requires strong Python coding skills, hands-on AWS experience, production ML deployment experience, and the ability to independently implement technical solutions. Candidates will also lead FinOps optimization, implement monitoring and alerting, ensure security and compliance, and mentor junior engineers.

San DiegoLast seen 24 days ago
$116,600 – $194,400 · Posted 26 days ago

As a Staff Data Engineer on Dexcom's Commercial Data Science & Revenue Operations team, you will design and operate the company's commercial data platform on AWS, building pipelines that ingest from 12+ vendor sources, transform data, and deliver analytics-ready datasets to data scientists, analysts, and 300+ field representatives. You will own AWS-native development using Lambda, Glue, CDK, Step Functions, and Redshift, with responsibility for data quality frameworks, alerting systems, warehouse schema design, and governance compliance.… The role requires hands-on proficiency in Python, SQL, AWS architectures, and CI/CD practices, along with experience building observability infrastructure and working with healthcare compliance requirements like HIPAA. You will also mentor team members and partner with business stakeholders to translate data capabilities into strategy.

San DiegoLast seen 24 days ago
$141,000 – $212,000 · Posted 27 days ago

Design, build, and operate scalable platform infrastructure across Azure, AWS, and private cloud environments using infrastructure-as-code, Kubernetes, and automation tooling. Develop reusable platform capabilities, CI/CD pipelines, and self-service infrastructure for engineering teams.… Own platform initiatives end-to-end from technical design through production operations, including deployment automation, configuration management, observability, and lifecycle management. Requires 7+ years of platform engineering or DevOps experience, strong hands-on proficiency with Terraform, Ansible, Python/Go, Kubernetes, and Linux systems administration.

San DiegoLast seen 25 days ago
$230,000 – $290,000 · Posted 1 month ago

Staff-level backend engineer responsible for setting technical strategy and owning search and recommendation systems at scale. You'll design and launch highly available distributed systems using Python or Kotlin, collaborate across product/design/analytics to balance technical sustainability with business impact, and lead a team through code review standards, mentorship, and operational excellence.… The role requires 8+ years of backend systems experience, 2+ years in search/recommendations, and proficiency with AWS, MySQL, Kubernetes, and modern architecture patterns.

San DiegoLast seen 1 month ago
$190,000 – $280,000 · Posted 1 month ago

Shield AI seeks an experienced Site Reliability Engineer to establish and mature the SRE function across Hivemind's cloud infrastructure and platform services. This hands-on leadership role involves defining reliability targets (SLIs/SLOs), improving observability and incident response practices, investigating complex failures, and driving adoption of reliability-first engineering practices.… The SRE Lead will mentor teams, develop operational automation tooling, and manage the multi-quarter SRE roadmap while working closely with Cloud Engineering and product teams to enhance system resilience and operational excellence.

San DiegoLast seen 1 month ago
Posted 1 month ago

Build and test agent behavior for an AI-driven service, including test case design and evaluation checks distinct from traditional software testing. Support data ingestion pipelines using AWS and BigData tools (PySpark, Airflow), collaborate with product and senior engineers to ship features, and contribute to production support (monitoring, alerting, incident response).… Requires 1–2 years of software engineering experience, a Master's in Computer Science or related field, proficiency in Python or another backend language, SQL/NoSQL fundamentals, basic REST API design, and hands-on experience with at least one LLM project (tool calling, chaining, retrieval, or multi-step coordination). Self-directed learner excited about AI tooling and focused on shipping end-to-end products.

San DiegoLast seen 1 month ago
$202,500 – $274,000 · Posted 1 month ago

Staff Data Engineer role at Credit Karma building petabyte-scale data infrastructure and streaming platforms on GCP. You will design and operate Kafka-based streaming infrastructure, build cloud-native data pipelines using Dataflow/Beam, Flink, and Spark, and create persistence frameworks across Spanner, MySQL, and BigQuery.… The role requires 7+ years backend/data systems experience in JVM languages (Scala/Java), expertise with high-throughput distributed systems, and deep knowledge of streaming platforms, Apache Beam, CDC patterns, encryption, and data governance frameworks. You'll also integrate AI/ML and generative AI technologies including LLMs, RAG, semantic search, and knowledge graphs into the data platform.

San DiegoLast seen 1 month ago
$92,400 – $148,800 · Posted 1 month ago

A Senior Site Reliability Engineer at SHEIN will own and operate mission-critical, large-scale distributed systems (Kubernetes, Kafka, Elasticsearch, Redis, APISIX, Nginx) running 24/7/365, participating in on-call rotations and driving fast incident response using AI-assisted log analysis and anomaly detection. The role requires strong software engineering expertise in Python or Go, deep Linux and networking knowledge, and hands-on experience with observability platforms (Prometheus, Grafana) and configuration management tools.… Responsibilities include designing resilient monitoring and alerting infrastructure, automating operational workflows to eliminate toil, capacity planning, and collaborating with global teams to improve system reliability and performance.

San DiegoLast seen 4 days ago
Posted 1 month ago

Own the complete release and deployment strategy for Elemynt's secure AI infrastructure platform, building systems that ensure every release is versioned, reproducible, observable, and safe to operate across enterprise and scientific computing environments. Design and implement deployment automation, CI/CD pipelines, artifact management, and operational visibility (logs, metrics, alerts) that work reliably in customer-managed and restricted environments.… Debug incidents, identify root causes, and systematize fixes into reusable automation. Partner with product and engineering teams to embed deployment readiness into the platform from the start.

San DiegoLast seen 11 days ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 1 month ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 1 month ago