← Back to results

observability jobs in San Diego

$149,800 – $262,200 · Posted 8 days ago

As a Staff Software Engineer in Data Platform Engineering, you will design, build, and operate cloud-native platform components and services, taking features from design through production. You'll lead well-scoped technical projects, collaborate with senior engineers on architecture alignment, and ensure reliability and scalability through testing and instrumentation. You'll spend most of your time hands-on building operators, controllers, automation, and platform services that other teams depend on, while mentoring junior engineers. The role requires 8+ years of software development experience (or equivalent with advanced degrees), 5+ years of hands-on Kubernetes experience, and proficiency with at least one major hyperscaler (AWS, Azure, GCP).

San DiegoLast seen 6 days ago
$130,814 – $130,814 · Posted 8 days ago

Design, implement, and maintain secure DevSecOps automation frameworks across on-premise and cloud infrastructure, integrating security controls into CI/CD pipelines for application and configuration deployment. Develop Infrastructure as Code using Terraform, Ansible, ARM/Bicep, PowerShell, Python, and Bash across Windows, Linux, VMware, Azure, and hybrid environments. Automate provisioning, patching, vulnerability remediation, and compliance checks aligned with FFIEC, NCUA, CIS, and NIST standards while partnering with Cybersecurity, Infrastructure, and Application Development teams. Support container/platform engineering, cloud identity governance, and promote DevSecOps and GitOps practices across technology teams.

San DiegoLast seen 6 days ago
Posted 9 days ago

Own the complete release and deployment strategy for Elemynt's secure AI infrastructure platform, building systems that ensure every release is versioned, reproducible, observable, and safe to operate across enterprise and scientific computing environments. Design and implement deployment automation, CI/CD pipelines, artifact management, and operational visibility (logs, metrics, alerts) that work reliably in customer-managed and restricted environments. Debug incidents, identify root causes, and systematize fixes into reusable automation. Partner with product and engineering teams to embed deployment readiness into the platform from the start.

San DiegoLast seen 7 days ago
Posted 9 days ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code. Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 7 days ago
Posted 9 days ago

Build the foundational agentic AI layer for a materials-science platform, including multi-model provider abstraction, agent orchestration with stateful checkpoints, retrieval systems, prompt versioning, and comprehensive tracing and evaluation frameworks. You'll design agents that plan and reason over tool calls in production, implement human-in-the-loop safety gates, and ensure all LLM behavior remains auditable and cost-tracked across customers' secure environments. The role demands deep production experience with agentic and LLM systems: async Python, structured outputs, memory and context management, multi-step workflow orchestration, and evaluation harnesses that catch regressions before deployment.

San DiegoLast seen 7 days ago
Posted 9 days ago

Ignite Digital seeks an AI Engineer to own the end-to-end machine learning lifecycle—from data preparation and model training through production deployment and monitoring. You will integrate AI models into frontend and backend systems, optimize performance for latency and cost, and collaborate with technical leads and customers to deliver Data and AI solutions. The role requires strong Python and cloud platform expertise (AWS, Azure, or GCP), software engineering fundamentals, and the ability to communicate technical opportunities and limitations across varying levels of technical experience.

San DiegoLast seen 7 days ago
$144,900 – $265,800 · Posted 9 days ago

Lead operational service delivery for enterprise Web Application Firewall (WAF), load balancing, DNS, certificates, and edge security platforms. Own end-to-end service performance, SLA/KPI metrics, incident escalation, and cross-team coordination across application, cloud, networking, and security teams. Manage and develop senior and staff-level engineers, drive operational governance, risk management, and continuous improvement. The role requires sufficient technical depth to guide platform designs, direct incident response, and make informed operational decisions, though not necessarily performing all configurations personally.

San DiegoLast seen 7 days ago
$140,000 – $180,000 · Posted 9 days ago

Lead an embedded Linux software team building mission-critical defense systems, splitting time between management and individual contribution. Design and implement C/C++ Linux daemons, REST APIs, and system services for aerial platforms and deployable manufacturing infrastructure. Manage team standups, architecture reviews, and cross-functional coordination with autonomy, avionics, and mechanical/electrical engineering. Engineer for real-time performance, security, and reliability while owning developer experience, CI/CD pipelines, and field operability.

San DiegoLast seen 7 days ago
$190,100 – $316,800 · Posted 10 days ago

Lead the architectural design of server-side, backend, and cloud platform capabilities supporting Dexcom's continuous glucose monitoring products. Collaborate across product, cybersecurity, firmware, data, and mobile teams to translate business and regulatory requirements into secure, scalable, and maintainable platform solutions. Establish architectural direction for distributed services, APIs, data pipelines, event-driven systems, and shared platform capabilities. Guide technology exploration, system integration, performance testing, and serve as an authority for platform architecture principles and modern backend development practices across the organization.

San DiegoLast seen 8 days ago
$147,000 – $163,000 · Posted 10 days ago

Lead Veracyte's Quality Engineering Platform team as Manager of Software Test Automation, defining quality strategy and building scalable automation frameworks that enable product teams to own software quality. Manage a team of SDETs, Automation Engineers, and Quality Platform Engineers to develop self-service testing platforms, reusable automation libraries, and AI-assisted quality tools supporting unit, integration, API, UI, end-to-end, regression, performance, security, and compliance testing. Establish modern engineering practices, quality standards, and reference architectures across the organization while fostering a culture of innovation and continuous improvement. Drive Veracyte's strategy for AI throughout the software testing lifecycle, including automated test generation, regression analysis, defect classification, and synthetic test data creation.

San DiegoLast seen 8 days ago
$128,900 – $219,100 · Posted 10 days ago

Design, build, and operate cloud-native platform features and services with hands-on engineering focused on Kubernetes, distributed systems, and infrastructure automation. Deliver well-defined work iteratively from implementation through production, contributing to reliability, scalability, and operability through testing, instrumentation, and on-call participation. Collaborate with senior engineers on platform standards while mentoring junior team members. Requires 5+ years of software development experience (or equivalent), production experience with Kubernetes, exposure to at least one major hyperscaler (AWS, Azure, GCP), and proficiency in Go or another systems language.

San DiegoLast seen 8 days ago
$130,814 – $163,517 · Posted 11 days ago

Design, implement, and maintain secure DevSecOps automation frameworks and CI/CD pipelines across hybrid infrastructure, cloud, and application platforms. Develop Infrastructure as Code using Terraform, Ansible, ARM/Bicep, PowerShell, and Python to automate provisioning, configuration management, patching, and vulnerability remediation across Windows, Linux, VMware, and Azure environments. Integrate security controls into deployment pipelines including vulnerability scanning, secrets detection, code scanning, and policy enforcement aligned with FFIEC, NCUA, CIS, and NIST standards. Partner across Infrastructure, Security, and Application Development teams to advance DevSecOps practices, container engineering, and secure cloud governance.

San DiegoLast seen 9 days ago
$125,000 – $160,000 · Posted 11 days ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response. Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 9 days ago
$108,400 – $127,500 · Posted 12 days ago

Join Singular Genomics' core software team as a generalist Software Engineer contributing to the sequencing analysis pipeline, operational tooling, and field support. You'll work across Python development, CI/CD infrastructure (Jenkins, Bitbucket), HPC job execution, and diagnostic troubleshooting of production failures on genomic sequencers. Responsibilities include pipeline bug fixes and features, build system maintenance, field triage, data handling logic, and developing lightweight web services and internal APIs. You need 2–5 years of professional experience, solid Python fundamentals, Linux/HPC familiarity, and the ability to debug and move fast across a large, real production codebase.

San DiegoLast seen 10 days ago
$125,000 – $160,000 · Posted 12 days ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes. Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 10 days ago
$108,000 – $180,000 · Posted 12 days ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper). Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 10 days ago
Posted 12 days ago

Lead end-to-end development of AI agents and autonomous engineering tools for Teradata's enterprise platform, managing multiple global engineering teams. Shape engineering strategy for agent platforms, LLM gateways, vector stores, observability frameworks, and low-code/no-code/pro-code agent frameworks. Drive architecture standardization, governance practices, operationalization pipelines, and ROI metrics while partnering with product, architecture, and AI platform teams. Requires 12+ years software engineering experience with 5+ years in technical leadership, hands-on expertise with LLMs, agent frameworks, distributed systems, cloud platforms (AWS, Azure, GCP), and modern CI/CD tools.

San DiegoLast seen 10 days ago