← Back to results

distributed-tracing jobs in San Diego

$152,216 – $152,216 · Posted 3 days ago

Cloud Engineering Manager leading enterprise cloud infrastructure teams at a credit union, requiring 9+ years of progressive cloud engineering experience with hands-on expertise in Azure, CI/CD, containerization, and infrastructure-as-code. The role combines deep technical work—architecture, troubleshooting, AKS management, database engineering, identity/access management, and cloud security—with team leadership, mentoring, and operational oversight of digital banking services.… Must demonstrate proficiency across Docker, Kubernetes, Terraform, monitoring (Datadog), and financial-industry compliance standards while driving DevOps automation and cost optimization.

San DiegoLast seen 2 days ago
$160,000 – $290,000 · Posted 10 days ago

Staff Engineer who will design, build, and operate the Forge Platform—a distributed-systems foundation serving autonomy, ML Ops, simulation, and application teams. You will own architecture and technical standards for workflow orchestration, asynchronous processing, event-driven systems, and long-running service workflows while remaining hands-on in implementation, production troubleshooting, and reliability improvement.… Required: strong production experience with distributed systems, cloud-native platforms, or backend infrastructure; fluency in Go and Python; deep understanding of failure handling, consistency, fault tolerance, and state management; and ability to turn recurring infrastructure needs into reusable platform capabilities. You will work across Kubernetes, service networking, observability, and data pipelines to enable downstream teams to move faster on mission-critical systems.

San DiegoLast seen 8 days ago
$160,000 – $290,000 · Posted 10 days ago

Staff Engineer responsible for designing, building, and operating the Forge Platform—a distributed-systems foundation serving autonomy, ML, simulation, and test teams. You will architect cloud-native backend infrastructure, establish reusable platform capabilities, define technical standards, and remain hands-on in production troubleshooting, performance analysis, and reliability improvement.… Required: deep expertise in production distributed systems, Go and Python, workflow orchestration, and the ability to translate complex architecture into usable interfaces across multiple downstream teams.

San DiegoLast seen 9 days ago
$123,000 – $185,000 · Posted 20 days ago

Staff Platform Engineer at Shield AI will design, build, and operate highly available infrastructure platforms supporting mission-critical systems across cloud and on-prem environments. Responsibilities include improving CI/CD pipelines, developer platform tooling, and automation frameworks; implementing scalable solutions for compute, networking, identity, and storage; and partnering with security teams to embed controls and compliance into platform architectures.… The role requires 8+ years of hands-on platform engineering, DevOps, or SRE experience with deep expertise in Infrastructure as Code (Terraform, Ansible), Linux systems, distributed systems, and scripting/programming languages (Python, Go, Bash, PowerShell).

San DiegoLast seen 18 days ago
Posted 21 days ago

Staff Platform Engineer at Shield AI responsible for designing, building, and operating highly available cloud and hybrid infrastructure platforms supporting mission-critical systems. The role spans infrastructure-as-code development, CI/CD pipeline enhancement, developer tooling, and platform automation to accelerate engineering productivity.… Requires 8+ years of platform engineering, DevOps, or SRE experience with deep hands-on expertise in cloud infrastructure, networking, identity, storage, and infrastructure automation across Linux environments. Expected to partner with security teams on compliance and DevSecOps practices, establish scalable platform patterns, and lead complex platform initiatives across large-scale engineering organizations.

San DiegoLast seen 15 days ago
$187,363 – $265,900 · Posted 1 month ago

Design and architect next-generation ML inference infrastructure for globally distributed, multi-tenant model serving with high availability, scalability, and cost efficiency. Lead the development of low-latency, high-throughput inference systems supporting computer vision and multimodal models (CNNs, segmentation, object detection) using Scala, Java, and Go.… Build large-scale distributed systems with reactive frameworks, integrate enterprise feature stores, and extend CI/CD pipelines (GitHub Prow, Pulumi) with automation and policy enforcement. Design advanced observability frameworks, optimize ML algorithms for performance, mentor engineers, and ensure MLOps and compliance standards (GDPR, SOC2) across the platform.

San DiegoLast seen 1 month ago
$120,001 – $160,000 · Posted 1 month ago

SAIC seeks a Cloud DevOps Engineer to design, build, and operate mission-critical cloud infrastructure on AWS and Kubernetes/OpenShift platforms in support of Naval Operational Architecture. The role combines infrastructure-as-code (CloudFormation, YAML), container orchestration, Linux/Windows server administration, networking, scripting (Bash, Python, PowerShell), and cybersecurity hardening (STIGs, ACAS).… You will work across cyber, networking, and development teams to deliver resilient, secure, and highly available cloud environments while maintaining compliance and system uptime.

San DiegoLast seen 1 month ago
$160,001 – $200,000 · Posted 1 month ago

Build and maintain AWS cloud infrastructure (EC2, VPCs, security groups, storage) using Infrastructure as Code with CloudFormation. Operate Kubernetes/OpenShift container platforms and manage application lifecycle.… Support web application stacks (React, Node.js, PostgreSQL) and implement authentication via OpenID Connect, SAML, and LDAP/AD. Monitor systems using CloudWatch and distributed tracing tools.

San DiegoLast seen 1 month ago
Posted 1 month ago

This Staff Platform Engineer role owns operational strategy and technical direction for enterprise GitLab (3,000+ users) and JFrog Artifactory platforms, leading cloud-native transformation and Kubernetes migrations. The position requires 7+ years of enterprise-scale platform operations, 5+ years each with self-managed GitLab and Artifactory in high-availability deployments, and 3+ years of deep Kubernetes expertise.… Responsibilities include designing multi-region active-active systems, establishing cloud-native best practices, mentoring team members, and influencing organizational standards for platform engineering and developer experience. Required expertise spans Linux kernel/networking, PostgreSQL high availability, AWS/Azure/GCP, Kubernetes operators, service mesh (Istio), infrastructure-as-code (Terraform), observability (Prometheus), and production automation in Python, Go, and Bash.

San DiegoLast seen 1 month ago