← Back to results

service-level-objectives jobs in San Diego

Posted 3 days ago

Lead enterprise data and AI architecture engagements at Deloitte, designing end-to-end Snowflake solutions for finance teams. You'll architect data pipelines, governance, ML/AI systems, and semantic models using Snowflake, Python, Snowpark, and Cortex AI.… The role combines technical leadership—guiding SQL/Python pipeline design, data quality, and MLOps—with client management, mentoring, and delivery accountability across planning, forecasting, close, and reporting use cases.

San DiegoLast seen 2 days ago
Posted 1 month ago

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production services running on AWS and Kubernetes. Responsibilities include designing highly available infrastructure with Terraform and managed AWS services, building CI/CD pipelines with GitHub Actions and Argo CD, defining SLOs and implementing observability with Prometheus and Grafana, leading incident response, and automating operational work with Python, Go, or Bash.… The role requires 3–8 years of hands-on SRE or DevOps experience, production Kubernetes expertise, strong AWS and infrastructure-as-code knowledge, and proficiency with observability and deployment strategies.

San DiegoLast seen 29 days ago
$138,400 – $173,000 · Posted 1 month ago

Sr. Site Reliability Engineer responsible for building and improving observability infrastructure, monitoring systems, and reliability patterns across AppFolio's Real Estate Platform.… You'll work with engineering teams to implement SLIs/SLOs, diagnose performance issues across the full stack, and manage infrastructure-as-code deployments on Kubernetes and AWS. Strong coding skills (Go, Ruby, or Python) and 5+ years of industry experience required; you'll be on-call and expected to help teams become self-sufficient in reliability practices.

San DiegoLast seen 24 days ago
$220,203 – $275,254 · Posted 1 month ago

Lead a team of cloud engineers building the backend platform that powers Brain Corp's global fleet of 30,000+ autonomous mobile robots. Own the platform's reliability, roadmap, and architecture as it scales to support fleet operations and customer-facing applications.… Balance hands-on technical leadership with people management—mentor engineers, drive hiring and onboarding, own production reliability and incident response, and partner with principal/staff engineers on distributed-systems architecture. Navigate tradeoffs between performance, cost, reliability, and scalability while shipping features that robots in the field depend on.

San DiegoLast seen 1 month ago
Posted 1 month ago

Senior Site Reliability Engineer responsible for designing and implementing Infrastructure as Code solutions, managing containerized and serverless architectures, and driving automation to reduce operational toil across large-scale production systems. The role requires defining Service Level Objectives (SLOs), leading incident response efforts, implementing monitoring and logging systems, and collaborating with development teams to ensure reliability and scalability.… Requires 5+ years of SRE or cloud engineering experience with deep expertise in Terraform, CloudFormation, Docker, Kubernetes, CI/CD pipelines, and cloud security practices. Must be a US citizen eligible to obtain and maintain an active Secret security clearance or above.

San DiegoLast seen 1 month ago
Posted 1 month ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code.… Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 15 days ago