← Back to results

opentelemetry jobs in San Diego

$150,000 – $200,000 · Posted 4 days ago

Cloud Engineer III will deploy, configure, administer, and troubleshoot cloud-based compute, storage, networking, and application infrastructure on AWS and Kubernetes. The role requires hands-on expertise in Linux systems administration, containerized applications (Docker/Kubernetes), cloud networking, IAM, security hardening, and observability platforms.… Responsibilities include supporting production Kubernetes clusters (EKS), implementing Infrastructure as Code (Terraform/CloudFormation), integrating CI/CD pipelines, automating administration with Python/Bash, and troubleshooting complex issues across cloud, container, and application layers. The position demands strong security practices, compliance knowledge for government environments, and collaboration with cross-functional technical and operational teams.

San DiegoLast seen 2 days ago
$150,000 – $225,000 · Posted 4 days ago

Build production AI features and workflows powered by large language models, including retrieval-augmented generation (RAG), agentic orchestration, and tool/function calling integrations. You will design and implement scalable microservices in Python or Java that connect foundation model APIs to business data and internal systems, with responsibility for AI quality, evaluation, observability, and safe deployment on AWS.… Collaborate with cross-functional teams to prototype solutions, measure outcomes, and move proven capabilities into production while maintaining reliability, security, and auditability.

San DiegoLast seen 2 days ago
Posted 11 days ago

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines.… The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.

San DiegoLast seen 9 days ago
Posted 11 days ago

The DevOps Engineer owns the design, deployment, and operation of production infrastructure at scale across AWS and Kubernetes, with responsibility for CI/CD pipelines, observability, and incident response. The role requires hands-on expertise in AWS services (EC2, EKS, VPC, IAM, S3, RDS), Kubernetes cluster operations, Terraform-based infrastructure-as-code, and monitoring tools like Prometheus and Grafana.… You will automate operational workflows with Python, Go, or Bash, lead incident investigations, and partner with software engineers and SREs to improve release velocity and system reliability. 3–8 years of DevOps, SRE, or platform engineering experience is expected.

San DiegoLast seen 9 days ago
$160,000 – $290,000 · Posted 14 days ago

Staff Engineer who will design, build, and operate the Forge Platform—a distributed-systems foundation serving autonomy, ML Ops, simulation, and application teams. You will own architecture and technical standards for workflow orchestration, asynchronous processing, event-driven systems, and long-running service workflows while remaining hands-on in implementation, production troubleshooting, and reliability improvement.… Required: strong production experience with distributed systems, cloud-native platforms, or backend infrastructure; fluency in Go and Python; deep understanding of failure handling, consistency, fault tolerance, and state management; and ability to turn recurring infrastructure needs into reusable platform capabilities. You will work across Kubernetes, service networking, observability, and data pipelines to enable downstream teams to move faster on mission-critical systems.

San DiegoLast seen 12 days ago
$160,000 – $290,000 · Posted 14 days ago

Staff Engineer responsible for designing, building, and operating the Forge Platform—a distributed-systems foundation serving autonomy, ML, simulation, and test teams. You will architect cloud-native backend infrastructure, establish reusable platform capabilities, define technical standards, and remain hands-on in production troubleshooting, performance analysis, and reliability improvement.… Required: deep expertise in production distributed systems, Go and Python, workflow orchestration, and the ability to translate complex architecture into usable interfaces across multiple downstream teams.

San DiegoLast seen 13 days ago
Posted 19 days ago

Lead Shield AI's Shared Services Engineering team, building and operating common APIs, SDKs, and customer-facing applications across the product portfolio. This is a hands-on technical management role requiring 7+ years of software engineering experience including team leadership, where you'll guide architecture and delivery of full-stack systems, establish secure and versioned interfaces for internal and external customers, and balance roadmap priorities across product needs and operational health.

San DiegoLast seen 17 days ago
$170,000 – $310,000 · Posted 27 days ago

Lead Shield AI's Shared Services Engineering team as a technical manager responsible for building, hiring, and coaching engineers who deliver cross-product infrastructure, SDKs, APIs, and customer-facing systems. Define and execute the team's roadmap while remaining hands-on in architecture and complex engineering decisions, balancing customer needs, operational health, and long-term technical investment.… Establish secure, reliable, and well-documented interfaces for internal and external users, and communicate plans and tradeoffs across engineering, product, and leadership.

San DiegoLast seen 25 days ago
Posted 27 days ago

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production services running on AWS and Kubernetes. Responsibilities include designing highly available infrastructure with Terraform and managed AWS services, building CI/CD pipelines with GitHub Actions and Argo CD, defining SLOs and implementing observability with Prometheus and Grafana, leading incident response, and automating operational work with Python, Go, or Bash.… The role requires 3–8 years of hands-on SRE or DevOps experience, production Kubernetes expertise, strong AWS and infrastructure-as-code knowledge, and proficiency with observability and deployment strategies.

San DiegoLast seen 25 days ago
$233,000 – $350,000 · Posted 1 month ago

Senior Staff Engineer responsible for designing and building an MLOps platform that supports distributed AI training, reinforcement learning, and foundation model development at scale. You will architect Kubernetes-native infrastructure for GPU workloads, design self-service AI development workflows, manage the data and model lifecycle, and lead platform distribution across cloud, on-premises, and air-gapped environments.… The role requires deep expertise in modern AI frameworks (PyTorch, Hugging Face Transformers), distributed systems, GPU scheduling, and cloud-native infrastructure, with a focus on enabling researchers and engineers to move from experimentation to production rapidly.

San DiegoLast seen 1 month ago
$140,000 – $185,000 · Posted 1 month ago

Senior Full Stack Engineer owning end-to-end manufacturing software platform development across API, frontend, data model, and background jobs. You will design and evolve core domain models (work orders, engineering changes, effectivity, defect records) and ship production-quality features that technicians, supervisors, and quality teams rely on daily.… The role requires deep expertise in TypeScript, React, Node.js backends, and Postgres with strong proficiency in schema design, migrations, and transactional consistency. You'll partner with product, design, and manufacturing stakeholders to translate operational needs into maintainable software while raising standards on testing and observability.

San DiegoLast seen 17 days ago
$150,000 – $225,000 · Posted 1 month ago

Build production AI features and services using foundation model APIs, retrieval-augmented generation (RAG), and agentic workflows. You will implement scalable microservices in Python or Java, develop evaluation and quality frameworks for AI systems, and collaborate across teams to integrate AI capabilities into enterprise workflows.… The role requires hands-on applied AI experience, strong backend or platform engineering skills, and familiarity with vector databases, LLM orchestration tools, and cloud infrastructure.

San DiegoLast seen 4 days ago
$158,400 – $237,600 · Posted 1 month ago

Design and build scalable, distributed data platform systems on AWS and Databricks, owning workspace architecture, governance, and lifecycle management. Implement infrastructure as code with Terraform, develop end-to-end data pipelines from ingestion to analytics and ML, and operate Amazon EKS clusters.… Define SLIs/SLOs for platform reliability, lead incident response and on-call rotations, implement security and compliance controls, and mentor engineers on platform standards. This senior staff role requires 8+ years building large-scale cloud platforms, 5+ years in Data Platform/DevOps/SRE roles, expert AWS knowledge, and deep hands-on experience with Databricks Lakehouse (Unity Catalog, Delta Lake, MLflow), Terraform, CI/CD pipelines, and Python/Bash scripting.

San DiegoLast seen 5 days ago
Posted 1 month ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code.… Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 11 days ago
Posted 1 month ago

Build and operate the data and ML infrastructure powering an AI platform for materials science, owning both sides: data pipelines that ingest and curate large-scale scientific output into training-ready formats, and model packaging, serving, monitoring, and CI/CD systems that move models safely from research to production across customer environments. You will design data ingestion and transformation workflows, implement validation and quality gates, package and version models with reproducible builds, run models through batch and online inference with safe rollout and rollback, monitor for drift and degradation, and build observability and internal tooling for engineering and science teams.… The role requires 6+ years shipping production software with deep expertise in data systems, ML infrastructure, containers, orchestration, and observability.

San DiegoLast seen 11 days ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 1 month ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 1 month ago