← Back to results

observability jobs in San Diego

$174,000 – $185,000 · Posted 1 month ago

Build and optimize a high-throughput compute stack for reliable, production-grade data ingestion, processing, storage, and delivery. The role focuses on making stateful, multi-threaded pipelines fast and observable under real-world constraints—profiling bottlenecks, hardening recovery behavior, and productizing machine-learning models (neural networks, tree-based, unsupervised) to meet performance, reliability, and quality standards.… You will work alongside algorithm and infrastructure engineers, using measurements to drive performance improvements and ensure production interfaces remain durable as the system evolves. Requires strong systems-level experience with concurrency, I/O, and production debugging in C++, Rust, CUDA, or C; a track record of shipping production software; and hands-on expertise in resilient data pipelines and GPU computing.

San DiegoLast seen 1 month ago
$150,000 – $180,000 · Posted 1 month ago

Senior Enterprise Applications Developer responsible for designing, delivering, and supporting complex enterprise software, integrations, data solutions, and automation across cloud and on-premises environments. Requires 7+ years of progressive experience with advanced Python proficiency, hands-on expertise in REST APIs, ETL/data pipelines, CI/CD practices, and AWS or Azure cloud-native development.… Must demonstrate ability to lead solution design from requirements through production support, with experience in workflow automation, AI/LLM integrations, and compliance frameworks (HIPAA, SOC 2, SOX). Strong emphasis on security, governance, architecture, and cross-functional collaboration with technical teams and business leaders.

San DiegoLast seen 1 month ago
$125,700 – $194,800 · Posted 1 month ago

As a Software Engineer (IC2) on ServiceNow's Data Platform team, you will design and develop scalable backend components for SQL generation into Trino, PostgreSQL, MariaDB, and Oracle databases, as well as manage table metadata and data access across multiple data sources. You'll own features from design through delivery, collaborate with product managers to translate requirements into well-architected solutions, and develop comprehensive test strategies covering functional, regression, integration, and performance aspects.… The role requires proficiency in Java backend development, strong knowledge of data structures, algorithms, and performance optimization, plus experience with relational databases, CI/CD pipelines, and automated testing frameworks. You'll participate in design and code reviews, troubleshoot complex systems, and foster engineering craftsmanship across the team.

San DiegoLast seen 1 month ago
Posted 1 month ago

This Staff AI Engineer role focuses on architecting and implementing secure, scalable frameworks for autonomous AI agents at enterprise scale. You'll lead the technical roadmap for agent management—from discovery and provisioning through task orchestration—while ensuring safety, compliance, and governance by design.… Working cross-functionally with product, platform, and research teams, you'll build foundational infrastructure for multi-agent collaboration, memory management, and extensibility, driving adoption across the enterprise. The role requires 8+ years in software architecture, backend systems, and Kubernetes/AI infrastructure, with deep expertise in LLM frameworks (LangChain, AutoGen, CrewAI), cloud platforms (AWS, Azure, GCP), and identity/access control patterns.

San DiegoLast seen 1 month ago
$211,000 – $297,000 · Posted 1 month ago

Principal Software Engineer role focused on designing highly complex enterprise-level web applications with deep expertise in microservices, serverless architectures, and cloud platforms. The engineer will serve as an architectural subject matter expert, providing guidance on system design, performance, scalability, and security across multiple teams at CoStar.… Requires 10+ years of hands-on experience with expert-level proficiency in C#, Java, Python, or JavaScript/TypeScript, plus significant hands-on experience designing and implementing AWS solutions. Responsibilities include creating architectural documentation, mentoring technical staff, staying current with emerging technologies, and advocating for security and observability best practices.

San DiegoLast seen 1 month ago
$233,000 – $350,000 · Posted 1 month ago

Senior Staff Engineer responsible for designing and building an MLOps platform that supports distributed AI training, reinforcement learning, and foundation model development at scale. You will architect Kubernetes-native infrastructure for GPU workloads, design self-service AI development workflows, manage the data and model lifecycle, and lead platform distribution across cloud, on-premises, and air-gapped environments.… The role requires deep expertise in modern AI frameworks (PyTorch, Hugging Face Transformers), distributed systems, GPU scheduling, and cloud-native infrastructure, with a focus on enabling researchers and engineers to move from experimentation to production rapidly.

San DiegoLast seen 1 month ago
$132,500 – $366,300 · Posted 1 month ago

Lead the design and engineering of enterprise-ready AI agents with retrieval, orchestration, policy-based routing, and lifecycle observability across multiple AI providers. Develop cloud-native abstraction layers, containerized microservices, and serverless architectures to deliver scalable agentic systems.… Tailor domain-specific workflows for finance, healthcare, and retail verticals through intelligent automation. Conduct design workshops, POCs, and code sessions with clients to shape data-driven agent workflows while defining metrics, evaluation harnesses, and best practices for agent accuracy, latency, safety, and cost-effectiveness.

San DiegoLast seen 1 month ago
$150,000 – $170,000 · Posted 1 month ago

Lead the design, governance, and technical implementation of an enterprise API and AI gateway platform serving internal and external consumers. Define solution standards, security controls, compliance protocols aligned to financial services regulations, and LLM integration patterns across Apigee X, MCP servers, and multiple AI providers.… Manage a team through API modernization and AI gateway buildout while actively contributing hands-on engineering on proxies, policies, and MCP integrations when needed.

San DiegoLast seen 1 month ago
$200,000 – $250,000 · Posted 1 month ago

Lead machine learning strategy and development for AppFolio's Leasing products, owning the ML roadmap and autonomous leasing agent architecture. Build evaluation frameworks, model quality infrastructure, and establish ML standards across the Leasing Engineering team while ensuring production-grade reliability, SLOs, and observability.… Translate research into shipped features by evaluating fine-tuning approaches, RAG patterns, and agentic systems; operate with production discipline on a SaaS platform serving real customer workflows.

San DiegoLast seen 14 days ago
$150,000 – $175,000 · Posted 1 month ago

As a Senior Platform Engineer at ReflexAI, you will design and maintain multi-cloud infrastructure across GCP and AWS using Terraform, Kubernetes, and GitHub, while building golden paths and guardrails that improve developer velocity. You'll own observability and monitoring with Datadog, strengthen security posture through IAM and container security best practices, and establish compliance workflows for SOC 2, HIPAA, and other industry standards.… This is a hands-on role requiring 5–7+ years of DevOps/SRE/platform engineering experience, deep cloud expertise, and CI/CD pipeline proficiency. You'll work cross-functionally to reduce operational toil and support an AI-driven product serving crisis support, healthcare, and enterprise customers.

San DiegoLast seen 1 month ago
$184,600 – $277,000 · Posted 1 month ago

Lead multiple engineering teams building social experiences on the PlayStation platform, translating product objectives into technical plans across client applications, APIs, distributed services, and cloud infrastructure. Establish standards for software design, testing, deployment, and operations while managing 24/7/365 service reliability, incident response, and observability.… Coach engineers across career stages, partner cross-functionally with Product, Design, Data, and Security leaders, and evaluate responsible AI integration across development, testing, and customer-facing capabilities.

San DiegoLast seen 1 month ago
$198,500 – $297,700 · Posted 1 month ago

Lead the design, development, and operation of Qualcomm's enterprise AI platform, supporting agentic AI, model lifecycle management, and multi-cloud ML/inference infrastructure. Own core platform services including identity/RBAC, secrets, service meshes, observability, vector stores, and model gateways across on-prem GPU clusters and managed cloud services.… Manage a ~10-engineer global team (platform, SRE, MLOps/LLMOps), drive incident response and continuous improvement, and partner with product and security teams on AI governance and high-impact use cases.

San DiegoLast seen 16 days ago
$142,100 – $213,100 · Posted 1 month ago

Staff Analytics Engineer responsible for designing and operationalizing agentic AI workflows, ML models, and Databricks applications at scale. Will build multi-step agent pipelines combining rules, ML models, and reasoning to solve complex business problems, then productionize them with monitoring, drift detection, and retraining strategies.… Requires 5+ years of hands-on ML engineering or data science with production system ownership, deep Python proficiency, strong traditional ML foundations, and proven Databricks expertise including notebook apps, dashboards, and ML pipelines. Will serve as technical authority, mentor peers, and influence architectural decisions across teams.

San DiegoLast seen 15 days ago
$171,900 – $300,800 · Posted 1 month ago

Lead a team of full-stack software engineers building ServiceNow's AI Experience Framework (AIUX) — a platform delivering AI-first, conversation-first user interfaces through modular Lit-based web components. You'll manage product development, set technical direction for framework-level APIs, guide architecture decisions across the request path, and represent AIUX in cross-org discussions.… The role requires deep expertise in JavaScript and Java/C++/C#/Go, 10+ years of relevant technology experience, 5+ years managing core engineering teams, and the ability to solve complex problems spanning AI integration, performance at scale, and multi-team coordination.

San DiegoLast seen 1 month ago
$115,300 – $160,100 · Posted 1 month ago

Senior AI Software Engineer responsible for designing, developing, and deploying LLM-powered applications and AI-assisted development workflows within a large-scale enterprise Linux environment (~6M+ LOC, primarily C/C++). Build high-performance GPU-based inference pipelines using vLLM and modern frameworks, develop agentic AI workflows with RAG and tool calling, and integrate LLMs with vector databases and enterprise systems into production.… Collaborate across software engineers, AI researchers, and platform teams to productionize AI services while optimizing for performance, latency, scalability, and operational efficiency. Requires 4+ years software development, strong Linux/Red Hat experience, advanced C/C++, Python, LLM/RAG expertise, and GitLab CI/CD workflows.

San DiegoLast seen 17 days ago
Posted 1 month ago

DevOps Engineer to build, maintain, and optimize AWS cloud infrastructure and CI/CD deployment systems, working with ECS, EKS, and EC2 workloads. The role requires hands-on experience with Infrastructure as Code (Terraform), containerization (Docker/Kubernetes), monitoring tools (Grafana), and CI/CD pipelines, with growing ownership of services and systems.… You'll troubleshoot infrastructure and application issues, collaborate with engineering teams on deployment workflows, and apply AI tooling to improve infrastructure automation and operational efficiency. The position reports to a DevOps Manager and is based in San Diego on a hybrid schedule.

San DiegoLast seen 16 days ago
$138,000 – $224,400 · Posted 1 month ago

Build and integrate oncology applications across laboratory systems, data pipelines, and enterprise platforms at Eli Lilly. Develop AI-enabled capabilities for data retrieval, extraction, and analysis while owning application reliability, user support, and lifecycle planning.… Deploy on Lilly cloud infrastructure using modern server-side frameworks (Python/Django/Flask preferred), front-end technologies (JavaScript/TypeScript), relational databases, and CI/CD pipelines. Mentor junior engineers and collaborate directly with scientific and operational stakeholders.

San DiegoLast seen 1 month ago
$73,700 – $128,780 · Posted 1 month ago

This role combines incident response, forensic analysis, and security operations responsibilities, requiring hands-on experience with cybersecurity tools, patch management, and security system optimization. The analyst will develop automation scripts (Python, PowerShell, Bash), manage firewalls and VPN systems, leverage network observability and security telemetry for threat detection, and support incident response workflows.… The position demands 2–5 years of progressive cybersecurity experience, intermediate certifications (GSEC, CEH, CySA+), and familiarity with NGFW technologies, cloud platforms (AWS, Azure), and security frameworks.

San DiegoLast seen 1 month ago
$183,500 – $232,500 · Posted 1 month ago

Senior Staff Software Engineer owning end-to-end design, architecture, and delivery of a scalable AWS-based cloud platform integrating backend services, frontend applications, mobile clients, and IoT hardware. You'll make high-stakes architectural decisions independently, write production code across Node.js, React, React Native, Python, Rust, and C++, design infrastructure-as-code with AWS CDK, and work with PostgreSQL at depth.… Operating in a flat, autonomous team with no management hierarchy, you'll set engineering standards, drive technical direction, and operate critical systems from design through production support.

San DiegoLast seen 1 month ago
Posted 1 month ago

Staff Automation Engineer responsible for developing, testing, and maintaining production automation systems, data pipelines, and workflow tools that support engineering organizations. The role requires strong Python and Linux expertise, SQL/database knowledge, and experience building maintainable backend systems and automation workflows.… Responsibilities include troubleshooting production issues, improving system reliability and monitoring, integrating data from multiple sources, and leveraging AI-assisted development tools to accelerate implementation and testing. This is a hands-on individual contributor role focused on execution, delivery, and knowledge transfer during an organizational transition.

San DiegoLast seen 1 month ago
$150,000 – $230,000 · Posted 1 month ago

Staff Engineer leading the design and operation of Shield AI's data platform to support autonomous systems, ML workflows, and edge-to-cloud data flows. You will architect distributed storage and compute solutions, build first-party integrations for heterogeneous data sources, and establish operational standards for reliability, security, and observability.… The role requires deep expertise in data modeling, API design, distributed systems, Kubernetes infrastructure, and the ability to guide technical direction across autonomy, ML, test, and infrastructure teams.

San DiegoLast seen 19 days ago
$128,900 – $219,100 · Posted 1 month ago

Senior Software Engineer on the Data Platform team will design, build, and operate cloud-native platform features and services with Kubernetes and hyperscaler expertise. The role requires 5+ years of software development experience (or equivalent), hands-on production software delivery, and proficiency in Go or another systems language.… You'll own features end-to-end from design through production, write tests and instrumentation, participate in on-call, and mentor junior engineers through code review and pairing.

San DiegoLast seen 1 month ago
$152,216 – $152,216 · Posted 1 month ago

Lead a cloud engineering team at a credit union while maintaining hands-on expertise in Azure cloud architecture, CI/CD pipelines, containerization, and infrastructure-as-code. The role requires 9+ years of progressive cloud engineering experience with deep technical knowledge of identity/access management (Okta, Entra ID), Kubernetes/AKS, database operations (MongoDB, MySQL, SQL Server), and financial-industry security standards.… You will balance mentorship and team leadership with active troubleshooting, architecture decisions, and driving operational excellence across cloud platforms and DevOps automation.

San DiegoLast seen 1 month ago
$138,000 – $224,400 · Posted 1 month ago

Software Engineer at Eli Lilly developing oncology applications that integrate laboratory systems, data pipelines, and enterprise platforms. The role requires 5+ years of software engineering experience and involves building AI-enabled capabilities, designing scalable APIs and web applications, deploying to AWS cloud infrastructure, and owning application reliability and lifecycle.… Key responsibilities include implementing data quality standards, contributing to CI/CD and containerization practices, mentoring junior engineers, and working directly with scientific and operational stakeholders.

San DiegoLast seen 1 month ago
$138,400 – $173,000 · Posted 1 month ago

Sr. Site Reliability Engineer responsible for building and improving observability infrastructure, monitoring systems, and reliability patterns across AppFolio's Real Estate Platform.… You'll work with engineering teams to implement SLIs/SLOs, diagnose performance issues across the full stack, and manage infrastructure-as-code deployments on Kubernetes and AWS. Strong coding skills (Go, Ruby, or Python) and 5+ years of industry experience required; you'll be on-call and expected to help teams become self-sufficient in reliability practices.

San DiegoLast seen 20 days ago