← Back to results

sre jobs in San Diego

$139,100 – $254,900 · Posted 1 day ago

This is a hands-on Solution Architect role for an AI-enabled Managed Services team, requiring deep technical expertise in designing end-to-end application architectures spanning UI, APIs, services, data, and AI components. The role combines hands-on coding and prototyping with architecture ownership, emphasizing active participation in codebases, AI tool evaluation, and guiding engineering teams through complex technical decisions.… Key responsibilities include designing reference architectures for AI patterns (RAG, vector search, prompt orchestration), defining non-functional requirements (performance, reliability, security), architecting CI/CD pipelines, and ensuring production readiness. The ideal candidate has 8+ years of software engineering experience, strong cloud-native and modern application architecture knowledge, practical AI feature development, and comfort working directly in code and infrastructure.

San DiegoLast seen today
$137,100 – $227,000 · Posted 7 days ago

Lead end-to-end AI platform architecture and infrastructure at enterprise scale, designing and implementing generative AI capabilities including LLMs, RAG systems, and agentic AI with a focus on model serving, inference optimization, and production deployment. Establish MLOps/LLMOps best practices, build automated CI/CD pipelines tailored for AI/ML applications, and architect vector database integration and cloud AI/ML services.… Mentor engineering teams, drive cross-functional collaboration between software, product, and data teams, and continuously evaluate and optimize AI platform performance, scalability, and reliability.

San DiegoLast seen 5 days ago
$153,000 – $170,000 · Posted 7 days ago

Platform Operations Engineer to build and operate healthcare infrastructure on AWS EKS, managing container environments, CI/CD pipelines via GitHub Actions, and observability stacks with Datadog and CloudWatch. Responsibilities include supporting reliability through SLI/SLO tracking, on-call rotations, and incident response; maintaining GitOps workflows with ArgoCD; and writing infrastructure automation in Python.… Requires 3+ years SRE/DevOps experience, hands-on AWS (VPC, IAM, EKS, RDS), Terraform or CloudFormation, Kubernetes, and familiarity with observability tools.

San DiegoLast seen 5 days ago
$190,000 – $280,000 · Posted 11 days ago

Senior Staff Lead Site Reliability Engineer to establish and mature SRE practices across Hivemind's cloud infrastructure and platform services. You will define reliability targets (SLIs/SLOs), build observability systems, lead incident response and root-cause analysis, mentor teams on reliability-first practices, and develop automation to reduce manual operational work.… The role requires 7+ years in SRE or infrastructure engineering, hands-on experience operating production services in AWS or equivalent cloud environments, expertise with containerized/distributed systems, infrastructure-as-code, and operational tooling development in Python or Go.

San DiegoLast seen 9 days ago
Posted 12 days ago

As a Software Engineer on the Platform Engineering team, you will architect and deliver cloud-native developer platforms and services that power ResMed's digital health ecosystem. You'll own the full lifecycle from design through production operations, building observability and reliability into services while participating in 24/7 on-call incident management.… The role requires strong fundamentals in at least one programming language (Java, Python, Go, TypeScript, JavaScript, React), hands-on experience with AWS and cloud-native technologies (Kubernetes, Docker, serverless), and a developer-first mindset focused on improving tooling and automation for internal engineering teams.

San DiegoLast seen 10 days ago
$175,000 – $195,000 · Posted 12 days ago

Senior or Staff Cloud Infrastructure Engineer responsible for designing and operating Crucible, a manufacturing-network software platform running across commercial cloud, government cloud, and on-premises environments. You will own infrastructure-as-code (Terraform, Helm, GitOps), Kubernetes production operations, cloud networking, identity and security controls, CI/CD pipelines, and observability tooling.… You'll work directly with software engineers to improve developer velocity, ensure production reliability, and support complex multi-environment deployments including GovCloud and air-gapped systems.

San DiegoLast seen 10 days ago
$149,800 – $262,200 · Posted 15 days ago

Staff SRE for ServiceNow's Government Community Cloud, providing 24/7 production support across a 3-shift team. The role combines software development, systems engineering, and networking to maintain reliability, scalability, and performance of federal infrastructure, with emphasis on automation, incident reduction, and MTTR optimization.… Requires 8+ years of related experience (or equivalent education trade-off), deep Linux knowledge, 2+ years DevOps/CI-CD and cloud experience, coding proficiency in Python/JavaScript/Ruby, database administration, and observability expertise at scale.

San DiegoLast seen 13 days ago
Posted 15 days ago

As a Staff Site Reliability Engineer at Altium (Renesas), you will ensure reliability, availability, and performance of large-scale SaaS cloud platforms through a combination of software engineering and systems administration. You will pioneer improvements in observability (logging, monitoring, APM), develop reliability frameworks, contribute to incident response and management, and drive automation and Infrastructure as Code initiatives across multiple regions.… The role requires 6+ years of SRE/DevOps experience in large-scale environments, 3+ years of software development (ideally .NET), strong knowledge of Kubernetes, AWS, microservices, and HA architecture, plus hands-on expertise with observability tools, CI/CD platforms, and IaC tools.

San DiegoLast seen 13 days ago
$230,000 – $290,000 · Posted 18 days ago

Lead a Reliability Platform Engineering team at Affirm responsible for building observability, risk management, and operational intelligence systems that help engineers manage production reliability at scale. You'll translate operational challenges into technical requirements, develop scalable reliability capabilities leveraging AI and automation, and drive alignment across Platform Engineering, SRE, Infrastructure, and product teams.… The role requires 7+ years of backend/full-stack engineering experience with 2+ years of engineering leadership, deep expertise with observability tools and SRE practices, and strong programming skills in Python, Kotlin, Java, or similar languages.

San DiegoLast seen 17 days ago
Posted 18 days ago

This is a Staff Site Reliability Engineer role focused on ensuring reliability, availability, and performance of Altium's large-scale SaaS cloud platforms. The role combines software engineering and systems administration, requiring 6+ years of SRE/DevOps experience and 3+ years of software development (preferably .NET).… Responsibilities include designing observability frameworks, automating operational tasks, managing incidents, implementing infrastructure-as-code, and collaborating with engineering teams on reliability best practices across AWS and Kubernetes environments.

San DiegoLast seen 16 days ago
Posted 19 days ago

Shield AI seeks an experienced SRE Lead to establish and mature reliability practices across Hivemind's cloud infrastructure and platform services. This hands-on technical role involves defining SLIs/SLOs, building monitoring and alerting systems, leading incident response, and driving root-cause analysis.… The SRE Lead will mentor teams, develop operational tooling in Python or Go, and partner with product and cloud engineering teams to embed reliability-first practices into system design and infrastructure provisioning.

San DiegoLast seen 17 days ago
$114,400 – $114,400 · Posted 19 days ago

Senior Systems Engineer responsible for administration and operations of a global cloud SaaS infrastructure serving millions of users. The role requires expert-level Linux and systems administration skills, with hands-on responsibility for availability, performance monitoring, and troubleshooting across web, database, and network layers in production environments.… Must have SRE/DevOps experience managing large-scale server infrastructure, proficiency with configuration management tools (Ansible, Puppet, Chef), monitoring platforms (Nagios, Splunk), and scripting languages (Python, Perl, or JavaScript). Knowledge of databases (MySQL, PostgreSQL, Oracle), load balancing (F5, NGINX), and networking technologies is essential.

San DiegoLast seen 17 days ago
Posted 26 days ago

Design and build platform engineering infrastructure, tools, and automation for cloud-native environments on AWS, using Infrastructure as Code (Terraform/CDK), CI/CD, and observability practices. Develop reusable modules, APIs, CLIs, and service templates that improve developer experience and enforce security, compliance, and cost controls.… Establish monitoring, logging, alerting, SLOs, and incident response practices; participate in 24x7 on-call rotation with a focus on production reliability and automation. Requires 7+ years in DevOps, SRE, platform engineering, or cloud operations with deep hands-on experience in multi-account AWS setups, GitOps, and observability platforms like Datadog or Prometheus.

San DiegoLast seen 24 days ago
$171,900 – $300,800 · Posted 28 days ago

Lead a team of engineering managers and software engineers building distributed agentic systems and AI-powered developer platforms at ServiceNow. You'll drive technical strategy for highly scalable, fault-tolerant systems; mentor teams; and architect foundational capabilities that improve developer experience for AI-first development.… This role requires 8+ years of software engineering experience with proven distributed systems expertise, 3+ years managing high-performing teams, and deep hands-on knowledge of Node.js, TypeScript, Kubernetes, and cloud-native infrastructure.

San DiegoLast seen 25 days ago
$190,000 – $280,000 · Posted 1 month ago

Shield AI seeks an experienced Site Reliability Engineer to establish and mature the SRE function across Hivemind's cloud infrastructure and platform services. This hands-on leadership role involves defining reliability targets (SLIs/SLOs), improving observability and incident response practices, investigating complex failures, and driving adoption of reliability-first engineering practices.… The SRE Lead will mentor teams, develop operational automation tooling, and manage the multi-quarter SRE roadmap while working closely with Cloud Engineering and product teams to enhance system resilience and operational excellence.

San DiegoLast seen 1 month ago
$137,100 – $227,000 · Posted 1 month ago

Lead end-to-end AI platform architecture and design for generative AI solutions, including LLMs, RAG systems, and agentic AI. Build MLOps/LLMOps pipelines with automated CI/CD, model serving infrastructure, and vector database integration on enterprise cloud platforms.… Establish best practices, mentor engineering teams, and drive seamless AI application integration into production systems while optimizing for scalability, security, and performance.

San DiegoLast seen 1 month ago
$150,000 – $175,000 · Posted 1 month ago

As a Senior Platform Engineer at ReflexAI, you will design and maintain multi-cloud infrastructure across GCP and AWS using Terraform, Kubernetes, and GitHub, while building golden paths and guardrails that improve developer velocity. You'll own observability and monitoring with Datadog, strengthen security posture through IAM and container security best practices, and establish compliance workflows for SOC 2, HIPAA, and other industry standards.… This is a hands-on role requiring 5–7+ years of DevOps/SRE/platform engineering experience, deep cloud expertise, and CI/CD pipeline proficiency. You'll work cross-functionally to reduce operational toil and support an AI-driven product serving crisis support, healthcare, and enterprise customers.

San DiegoLast seen 1 month ago
Posted 1 month ago

DevOps Engineer to build, maintain, and optimize AWS cloud infrastructure and CI/CD deployment systems, working with ECS, EKS, and EC2 workloads. The role requires hands-on experience with Infrastructure as Code (Terraform), containerization (Docker/Kubernetes), monitoring tools (Grafana), and CI/CD pipelines, with growing ownership of services and systems.… You'll troubleshoot infrastructure and application issues, collaborate with engineering teams on deployment workflows, and apply AI tooling to improve infrastructure automation and operational efficiency. The position reports to a DevOps Manager and is based in San Diego on a hybrid schedule.

San DiegoLast seen 16 days ago
$150,000 – $200,000 · Posted 1 month ago

Staff-level database engineer responsible for managing production database environments across MySQL, ClickHouse, and MongoDB at scale. The role demands independent execution, proactive monitoring, performance optimization through query tuning and indexing, and 24/7 on-call incident response.… Requires 4+ years of MySQL and cloud platform experience (Azure/AWS), strong Linux expertise, and scripting proficiency (Golang preferred); nice-to-have skills include AI tooling integration (LangChain, LlamaIndex), vector databases, and RAG for automating database operations.

San DiegoLast seen 17 days ago
$188,500 – $255,000 · Posted 1 month ago

Staff Software Engineer to architect and lead Intuit's next-generation One Logging system, designing high-scale, mission-critical logging pipelines for on-premise and public cloud (AWS, GCP) deployments. Own end-to-end platform evolution from log ingestion through routing, storage, and cost optimization, including components like FELS, Kinesis/CloudWatch integrations, and Log Router.… Lead modernization of legacy tooling, design edge collection agents (Fluent Bit, OIL sidecar), and drive automation initiatives using MCP Server capabilities. Provide technical leadership, mentorship, and cross-functional collaboration to align logging platform capabilities across the organization.

San DiegoLast seen 1 month ago
$205,489 – $244,000 · Posted 1 month ago

Senior Staff Software Engineer leading teams to improve engineering efficiency and product quality in software development for 5G and RAN optimization. Designs and implements scalable software solutions, manages software architecture to meet performance and scalability demands of cellular networks, and serves as a tech lead owning project outcomes.… Works with site reliability engineers to translate telecom requirements into robust code, applies distributed systems and AI knowledge, and mentors development teams through deep technical expertise. Requires 6+ years of experience with a Master's degree or 8+ years with a Bachelor's degree in Computer Science, Computer Engineering, or Electrical Engineering.

San DiegoLast seen 1 month ago
$177,300 – $265,900 · Posted 1 month ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues.… Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 9 days ago
$86,400 – $138,000 · Posted 1 month ago

The Database Engineer will maintain, monitor, and optimize large-scale production MySQL database environments in distributed, mission-critical settings. Responsibilities include performance tuning, query and index optimization, database migrations and upgrades, capacity planning, and 24x7 on-call support.… The role requires designing reliable database architectures, automating operational tasks, and collaborating with development teams on database design guidance. Candidates must have strong fundamentals in distributed database topologies, Linux administration, scripting proficiency (Go/Python/Shell), and experience with cloud platforms such as Azure.

San DiegoLast seen 7 days ago
$190,100 – $316,800 · Posted 1 month ago

Lead the architectural design of server-side, backend, and cloud platform capabilities supporting Dexcom's continuous glucose monitoring products. Collaborate across product, cybersecurity, firmware, data, and mobile teams to translate business and regulatory requirements into secure, scalable, and maintainable platform solutions.… Establish architectural direction for distributed services, APIs, data pipelines, event-driven systems, and shared platform capabilities. Guide technology exploration, system integration, performance testing, and serve as an authority for platform architecture principles and modern backend development practices across the organization.

San DiegoLast seen 1 month ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 1 month ago