← Back to results

tracing jobs in San Diego

$149,800 – $262,200 · Posted 9 days ago

ServiceNow seeks a Staff Software Engineer (IC4) for their Data Platform Engineering organization to design, build, and operate cloud-native platform components and services on Kubernetes at scale. You will own well-scoped projects end-to-end, from design through production, collaborating with senior engineers and mentoring junior team members.… The role requires 8+ years of software development experience, 5+ years hands-on with Kubernetes, expertise in at least one major hyperscaler (AWS, Azure, GCP), and strong Go programming skills. You will focus on building operators, controllers, automation, and platform services that other teams depend on, with emphasis on reliability, scalability, operability, and infrastructure-as-code practices.

San DiegoLast seen 7 days ago
$190,000 – $280,000 · Posted 11 days ago

Senior Staff Lead Site Reliability Engineer to establish and mature SRE practices across Hivemind's cloud infrastructure and platform services. You will define reliability targets (SLIs/SLOs), build observability systems, lead incident response and root-cause analysis, mentor teams on reliability-first practices, and develop automation to reduce manual operational work.… The role requires 7+ years in SRE or infrastructure engineering, hands-on experience operating production services in AWS or equivalent cloud environments, expertise with containerized/distributed systems, infrastructure-as-code, and operational tooling development in Python or Go.

San DiegoLast seen 9 days ago
Posted 19 days ago

Shield AI seeks an experienced SRE Lead to establish and mature reliability practices across Hivemind's cloud infrastructure and platform services. This hands-on technical role involves defining SLIs/SLOs, building monitoring and alerting systems, leading incident response, and driving root-cause analysis.… The SRE Lead will mentor teams, develop operational tooling in Python or Go, and partner with product and cloud engineering teams to embed reliability-first practices into system design and infrastructure provisioning.

San DiegoLast seen 17 days ago
$150,100 – $225,100 · Posted 23 days ago

Software Engineer II responsible for operating and improving large-scale SQL, NoSQL, streaming, and caching platforms (Cassandra, Aurora, Aerospike, Kafka, Redis, DynamoDB, ElastiCache) with a focus on reliability, automation, and observability. Build self-service developer experiences using Terraform and infrastructure-as-code, participate in on-call incident response, and develop APIs and automated tests for infrastructure tooling.… Requires 3+ years in software engineering, database reliability engineering, SRE, or platform engineering, with hands-on Go experience, Terraform proficiency, AWS/GCP knowledge, and Kubernetes familiarity.

San DiegoLast seen 21 days ago
$150,100 – $225,100 · Posted 24 days ago

Software Engineer II responsible for operating and improving large-scale data platforms (SQL, NoSQL, streaming, caching) with a focus on reliability, automation, and self-service infrastructure. The role involves building observability tooling, participating in on-call incident response, troubleshooting database and infrastructure issues, and collaborating across teams to deliver reliable stateful services.… Required skills include Go, Terraform, API design, Kubernetes, and hands-on experience with databases like Cassandra, Aurora, Kafka, Redis, and DynamoDB on AWS or GCP.

San DiegoLast seen 22 days ago
$150,100 – $225,100 · Posted 24 days ago

Software Engineer II responsible for operating and improving large-scale data platforms (SQL, NoSQL, streaming, caching) with a focus on reliability, automation, and developer experience. The role involves building self-service infrastructure tooling, contributing to observability and SLO definitions, participating in on-call rotations, and automating infrastructure provisioning using Terraform.… Requires 3+ years of software or infrastructure engineering experience, production Go development, hands-on Terraform expertise, and deep knowledge of distributed systems concepts and operational practices.

San DiegoLast seen 23 days ago
$134,800 – $202,200 · Posted 27 days ago

Qualcomm Data Center seeks a Staff Linux Software Engineer to lead development of host-side system software, drivers, and libraries for next-generation datacenter inference accelerators. The role spans driver development, test infrastructure, performance optimization, and security enablement, with emphasis on production-ready Linux-based systems.… You will own complex subsystems end-to-end, drive architectural decisions, and provide technical leadership across software, hardware, and platform teams. Core responsibilities include designing Linux user-space and kernel-adjacent drivers in modern C++, architecting CI/test infrastructure, optimizing AI inference performance, and collaborating with hardware, firmware, and cloud teams.

San DiegoLast seen 26 days ago
$190,000 – $280,000 · Posted 1 month ago

Shield AI seeks an experienced Site Reliability Engineer to establish and mature the SRE function across Hivemind's cloud infrastructure and platform services. This hands-on leadership role involves defining reliability targets (SLIs/SLOs), improving observability and incident response practices, investigating complex failures, and driving adoption of reliability-first engineering practices.… The SRE Lead will mentor teams, develop operational automation tooling, and manage the multi-quarter SRE roadmap while working closely with Cloud Engineering and product teams to enhance system resilience and operational excellence.

San DiegoLast seen 1 month ago
$198,500 – $297,700 · Posted 1 month ago

Lead the design, development, and operation of Qualcomm's enterprise AI platform, supporting agentic AI, model lifecycle management, and multi-cloud ML/inference infrastructure. Own core platform services including identity/RBAC, secrets, service meshes, observability, vector stores, and model gateways across on-prem GPU clusters and managed cloud services.… Manage a ~10-engineer global team (platform, SRE, MLOps/LLMOps), drive incident response and continuous improvement, and partner with product and security teams on AI governance and high-impact use cases.

San DiegoLast seen 16 days ago
$149,800 – $262,200 · Posted 1 month ago

Staff Software Engineer role designing and implementing cloud-native platform components and services, with primary focus on Kubernetes workloads, distributed systems, and infrastructure automation. Requires 5+ years hands-on Kubernetes experience, proficiency in Go or systems languages, and demonstrated expertise building operators, controllers, and platform services.… Will lead well-scoped projects from design through production, contribute to reliability and scalability, mentor junior engineers, and participate in on-call operations.

San DiegoLast seen 1 month ago
$177,300 – $265,900 · Posted 1 month ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues.… Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 9 days ago
$149,800 – $262,200 · Posted 1 month ago

As a Staff Software Engineer in Data Platform Engineering, you will design, build, and operate cloud-native platform components and services, taking features from design through production. You'll lead well-scoped technical projects, collaborate with senior engineers on architecture alignment, and ensure reliability and scalability through testing and instrumentation.… You'll spend most of your time hands-on building operators, controllers, automation, and platform services that other teams depend on, while mentoring junior engineers. The role requires 8+ years of software development experience (or equivalent with advanced degrees), 5+ years of hands-on Kubernetes experience, and proficiency with at least one major hyperscaler (AWS, Azure, GCP).

San DiegoLast seen 1 month ago
Posted 1 month ago

Design, build, and ship LLM-powered capabilities end to end—from prototyping and fine-tuning models to deploying production agents and retrieval systems. Own the full stack: prompt and context engineering, multi-step agent design with tool calling, RAG systems (embeddings, chunking, hybrid search, reranking), fine-tuning on multi-GPU with LoRA/QLoRA, evaluation and tracing infrastructure, and clean APIs for other engineers.… Work in a secure, distributed environment where the platform runs on customer compute, cloud, or hybrid setups, requiring expertise with both commercial and self-hosted models.

San DiegoLast seen 11 days ago
Posted 1 month ago

Build the foundational agentic AI layer for a materials-science platform, including multi-model provider abstraction, agent orchestration with stateful checkpoints, retrieval systems, prompt versioning, and comprehensive tracing and evaluation frameworks. You'll design agents that plan and reason over tool calls in production, implement human-in-the-loop safety gates, and ensure all LLM behavior remains auditable and cost-tracked across customers' secure environments.… The role demands deep production experience with agentic and LLM systems: async Python, structured outputs, memory and context management, multi-step workflow orchestration, and evaluation harnesses that catch regressions before deployment.

San DiegoLast seen 11 days ago
$140,000 – $180,000 · Posted 1 month ago

Lead a small embedded Linux software team building mission-critical aerial systems and autonomous platforms for defense applications. You'll split your time between team management (standups, architecture reviews, task prioritization) and hands-on C/C++ development of Linux daemons, REST APIs, and system services.… Design for real-time performance, security hardening, and field operability while collaborating across autonomy, avionics, mechanical, and electrical teams. Own the full software lifecycle from architecture and CI/CD pipelines to documentation and field diagnostics.

San DiegoLast seen 1 month ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 1 month ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 1 month ago