← Back to results

elk jobs in San Diego

Posted 7 days ago

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines.… The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.

San DiegoLast seen 5 days ago
Posted 8 days ago

As a Software Engineer on the Platform Engineering team, you will architect and deliver cloud-native developer platforms and services that power ResMed's digital health ecosystem. You'll own the full lifecycle from design through production operations, building observability and reliability into services while participating in 24/7 on-call incident management.… The role requires strong fundamentals in at least one programming language (Java, Python, Go, TypeScript, JavaScript, React), hands-on experience with AWS and cloud-native technologies (Kubernetes, Docker, serverless), and a developer-first mindset focused on improving tooling and automation for internal engineering teams.

San DiegoLast seen 6 days ago
Posted 11 days ago

As a Staff Site Reliability Engineer at Altium (Renesas), you will ensure reliability, availability, and performance of large-scale SaaS cloud platforms through a combination of software engineering and systems administration. You will pioneer improvements in observability (logging, monitoring, APM), develop reliability frameworks, contribute to incident response and management, and drive automation and Infrastructure as Code initiatives across multiple regions.… The role requires 6+ years of SRE/DevOps experience in large-scale environments, 3+ years of software development (ideally .NET), strong knowledge of Kubernetes, AWS, microservices, and HA architecture, plus hands-on expertise with observability tools, CI/CD platforms, and IaC tools.

San DiegoLast seen 9 days ago
Posted 14 days ago

This is a Staff Site Reliability Engineer role focused on ensuring reliability, availability, and performance of Altium's large-scale SaaS cloud platforms. The role combines software engineering and systems administration, requiring 6+ years of SRE/DevOps experience and 3+ years of software development (preferably .NET).… Responsibilities include designing observability frameworks, automating operational tasks, managing incidents, implementing infrastructure-as-code, and collaborating with engineering teams on reliability best practices across AWS and Kubernetes environments.

San DiegoLast seen 12 days ago
$100,300 – $135,700 · Posted 1 month ago

Lead Information Assurance and security engineering for Navy tactical networks (CANES), managing end-to-end Risk Management Framework (RMF) authorization, hardening virtualized and cloud network stacks, and ensuring compliance across diverse naval environments. Develop and maintain security authorization artifacts (SSP, SAP/SAR, POA&M, ATO packages in eMASS), apply DISA STIGs and vulnerability management tools (ACAS, NESSUS, SCAP), and provide IA guidance for application integration aligned with DoDAF and Net-Ready KPP baselines.… Support developmental testing, audit response, and continuous monitoring while coordinating with PEO C4I, NIWC Pacific, and Fleet stakeholders. Mentor junior engineers and champion DevSecOps and STIG automation practices.

San DiegoLast seen 1 month ago