← Back to results

opentelemetry jobs in San Diego

$150,000 – $225,000 · Posted 3 days ago

Build production AI features and services using foundation model APIs, retrieval-augmented generation (RAG), and agentic workflows. You will implement scalable microservices in Python or Java, develop evaluation and quality frameworks for AI systems, and collaborate across teams to integrate AI capabilities into enterprise workflows. The role requires hands-on applied AI experience, strong backend or platform engineering skills, and familiarity with vector databases, LLM orchestration tools, and cloud infrastructure.

San DiegoLast seen 1 day ago
$158,400 – $237,600 · Posted 4 days ago

Design and build scalable, distributed data platform systems on AWS and Databricks, owning workspace architecture, governance, and lifecycle management. Implement infrastructure as code with Terraform, develop end-to-end data pipelines from ingestion to analytics and ML, and operate Amazon EKS clusters. Define SLIs/SLOs for platform reliability, lead incident response and on-call rotations, implement security and compliance controls, and mentor engineers on platform standards. This senior staff role requires 8+ years building large-scale cloud platforms, 5+ years in Data Platform/DevOps/SRE roles, expert AWS knowledge, and deep hands-on experience with Databricks Lakehouse (Unity Catalog, Delta Lake, MLflow), Terraform, CI/CD pipelines, and Python/Bash scripting.

San DiegoLast seen 2 days ago
Posted 9 days ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code. Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 7 days ago
Posted 9 days ago

Build and operate the data and ML infrastructure powering an AI platform for materials science, owning both sides: data pipelines that ingest and curate large-scale scientific output into training-ready formats, and model packaging, serving, monitoring, and CI/CD systems that move models safely from research to production across customer environments. You will design data ingestion and transformation workflows, implement validation and quality gates, package and version models with reproducible builds, run models through batch and online inference with safe rollout and rollback, monitor for drift and degradation, and build observability and internal tooling for engineering and science teams. The role requires 6+ years shipping production software with deep expertise in data systems, ML infrastructure, containers, orchestration, and observability.

San DiegoLast seen 7 days ago
$125,000 – $160,000 · Posted 11 days ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response. Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 9 days ago
$125,000 – $160,000 · Posted 12 days ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes. Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 10 days ago