← Back to results

observability jobs in San Diego

Posted 14 days ago

This Staff AI Engineer role focuses on architecting and implementing secure, scalable frameworks for autonomous AI agents at enterprise scale. The engineer will lead the technical roadmap for agent management—from provisioning and orchestration to compliance and retirement—while ensuring AI agents are secure by design with built-in visibility and guardrails.… The role requires 8+ years of software architecture and backend systems experience, deep knowledge of autonomous agent frameworks (LangChain, AutoGen, CrewAI), and expertise in cloud infrastructure (AWS, Azure, GCP) and Kubernetes. You'll collaborate cross-functionally with researchers, platform engineers, and product teams to bring agentic AI capabilities to production while establishing best practices for agent safety, observability, and ethical deployment.

San DiegoLast seen 13 days ago
$130,000 – $150,000 · Posted 14 days ago

Lead hands-on modernization of legacy SSIS/SSRS ETL workloads into cloud-native Lakehouse pipelines using Databricks, dbt, Fivetran, and Airflow. Design and enforce standardized data frameworks (ingestion, transformation, curation, consumption), drive CI/CD and observability practices, and mentor engineers on modern data engineering patterns.… Evaluate emerging technologies (Delta Live Tables, Iceberg, streaming, AI-driven observability) through POCs and POVs, and leverage AI-assisted tools (Databricks Assistant, GitHub Copilot, Cursor AI) to accelerate development and reduce technical debt. Optimize Spark workloads and orchestration across Azure and GCP environments while partnering with architects, platform engineers, and business stakeholders.

San DiegoLast seen 13 days ago
$128,900 – $219,100 · Posted 15 days ago

Design, build, test, and operate production ML and LLM-powered components for cybersecurity applications at scale. You will turn ambiguous requirements into working code, partner across product and security teams, and apply AI safety and guardrail practices.… The role requires solid software engineering fundamentals, hands-on experience building production ML/LLM applications (RAG, embeddings, agents), and the ability to take prototypes to reliable, maintainable systems. You'll work with distributed systems, cloud-native technologies, and graph/data plumbing while contributing to code reviews and raising team quality standards.

San DiegoLast seen 13 days ago
$149,800 – $262,200 · Posted 15 days ago

Staff SRE for ServiceNow's Government Community Cloud, providing 24/7 production support across a 3-shift team. The role combines software development, systems engineering, and networking to maintain reliability, scalability, and performance of federal infrastructure, with emphasis on automation, incident reduction, and MTTR optimization.… Requires 8+ years of related experience (or equivalent education trade-off), deep Linux knowledge, 2+ years DevOps/CI-CD and cloud experience, coding proficiency in Python/JavaScript/Ruby, database administration, and observability expertise at scale.

San DiegoLast seen 13 days ago
$199,000 – $258,000 · Posted 15 days ago

Sr. Manager overseeing a data engineering team with responsibility for hiring, coaching, and performance management while remaining hands-on with architecture, design reviews, and technical decisions.… The role requires designing and operating production data pipelines and lakehouse platforms using AWS (S3, Glue, Athena, Lambda) and Snowflake, with hands-on proficiency in SQL and Python. Responsibilities include establishing data governance, compliance, security controls, CI/CD practices, and collaborating across engineering, analytics, product, and clinical teams to translate requirements into platform capabilities. The manager must balance delivery priorities, technical debt reduction, operational excellence, and emerging-technology evaluation while driving outcomes that affect broader business objectives.

San DiegoLast seen 13 days ago
Posted 15 days ago

Saronic Technologies seeks experienced Systems Software Engineers to develop software for autonomous maritime systems, working across the full stack from low-level hardware interfaces and embedded systems to backend services and infrastructure. The role requires strong software engineering fundamentals and hands-on experience with cyber-physical systems, sensor/actuator integration, distributed services, and mission-critical reliability.… Candidates should have 5+ years of professional experience with C++, Python, Rust, or Go, Linux proficiency, and a track record building systems that interact with real-world hardware in robotics, autonomy, aerospace, or similar domains.

San DiegoLast seen 13 days ago
Posted 15 days ago

As a Staff Site Reliability Engineer at Altium (Renesas), you will ensure reliability, availability, and performance of large-scale SaaS cloud platforms through a combination of software engineering and systems administration. You will pioneer improvements in observability (logging, monitoring, APM), develop reliability frameworks, contribute to incident response and management, and drive automation and Infrastructure as Code initiatives across multiple regions.… The role requires 6+ years of SRE/DevOps experience in large-scale environments, 3+ years of software development (ideally .NET), strong knowledge of Kubernetes, AWS, microservices, and HA architecture, plus hands-on expertise with observability tools, CI/CD platforms, and IaC tools.

San DiegoLast seen 13 days ago
$138,000 – $224,400 · Posted 15 days ago

Software engineer responsible for building and integrating applications across laboratory systems, data pipelines, and enterprise platforms in an oncology context. The role emphasizes full-stack development (Python/Django/Flask backend, JavaScript/TypeScript frontend), cloud deployment on AWS, CI/CD and containerization practices, and increasingly AI-enabled capabilities such as LLM-powered data access and agentic workflows.… Requires 4+ years of professional software engineering experience, strong collaboration with scientific and operational stakeholders, and ownership of application reliability, documentation, and lifecycle planning.

San DiegoLast seen 13 days ago
$230,000 – $290,000 · Posted 18 days ago

Lead a Reliability Platform Engineering team at Affirm responsible for building observability, risk management, and operational intelligence systems that help engineers manage production reliability at scale. You'll translate operational challenges into technical requirements, develop scalable reliability capabilities leveraging AI and automation, and drive alignment across Platform Engineering, SRE, Infrastructure, and product teams.… The role requires 7+ years of backend/full-stack engineering experience with 2+ years of engineering leadership, deep expertise with observability tools and SRE practices, and strong programming skills in Python, Kotlin, Java, or similar languages.

San DiegoLast seen 17 days ago
Posted 18 days ago

This is a Staff Site Reliability Engineer role focused on ensuring reliability, availability, and performance of Altium's large-scale SaaS cloud platforms. The role combines software engineering and systems administration, requiring 6+ years of SRE/DevOps experience and 3+ years of software development (preferably .NET).… Responsibilities include designing observability frameworks, automating operational tasks, managing incidents, implementing infrastructure-as-code, and collaborating with engineering teams on reliability best practices across AWS and Kubernetes environments.

San DiegoLast seen 16 days ago
$138,000 – $224,400 · Posted 18 days ago

Software engineer developing oncology applications that integrate laboratory systems, data pipelines, and enterprise platforms. The role requires building AI-enabled capabilities (information retrieval, NLP extraction, summarization), deploying on Lilly cloud infrastructure, and owning application reliability and lifecycle.… Requires 4+ years of software engineering experience with BS/MS in computer science or related field; strong Python or modern server-side language skills, web frameworks (Django, Flask, Rails), front-end development, cloud deployment (AWS preferred), CI/CD, containerization, and familiarity with LLM APIs and agentic patterns are preferred.

San DiegoLast seen 16 days ago
Posted 19 days ago

Shield AI seeks an experienced SRE Lead to establish and mature reliability practices across Hivemind's cloud infrastructure and platform services. This hands-on technical role involves defining SLIs/SLOs, building monitoring and alerting systems, leading incident response, and driving root-cause analysis.… The SRE Lead will mentor teams, develop operational tooling in Python or Go, and partner with product and cloud engineering teams to embed reliability-first practices into system design and infrastructure provisioning.

San DiegoLast seen 17 days ago
$105,500 – $168,800 · Posted 19 days ago

Senior Software Engineer on the Pyxis ES Development team, responsible for designing and implementing scalable enterprise software solutions in microservices-based architectures. The role requires 5+ years of coding experience in languages like Python, Java, Go, C, or C++, with hands-on expertise in modern web technologies (React, Node.js, Spring, Ruby), containerization, messaging platforms (Kafka, RabbitMQ), and relational databases.… You will conduct software testing and evaluation, provide technical documentation, participate in design reviews, and work collaboratively with cross-functional teams in Agile environments. Strong analytical and problem-solving abilities, plus experience with cloud platforms and infrastructure-as-code tools, are essential.

San DiegoLast seen 17 days ago
Posted 19 days ago

Lead engineering and data science roadmap for Teradata's AI Platform, overseeing architecture, LLM integrations, RAG pipelines, vector store implementations, and data science experimentation frameworks. Partner with product, research, and engineering teams to define requirements for predictive modeling, multi-agent collaboration, and tool orchestration.… Manage and mentor a team of backend engineers, data scientists, and AI platform specialists while driving technical excellence, cloud-native practices, and platform scalability.

San DiegoLast seen 17 days ago
Posted 19 days ago

Sr. Staff IT Enterprise Architect will define integration standards and architectural decision frameworks across AppFolio's platforms (MuleSoft, AWS, Workato), promoting API-first and event-driven design.… The role partners with engineering teams to advance observability, reliability, and modern integration practices, while advising on AI-enabled capabilities and emerging technologies. Requires 10+ years of software engineering or integration architecture leadership with 8+ years designing enterprise integration solutions, hands-on platform experience, and demonstrated ability to influence across cross-functional teams without direct authority.

San DiegoLast seen 17 days ago
$200,000 – $250,000 · Posted 20 days ago

AppFolio is hiring a Staff Machine Learning Engineer to design, build, and operate their ML platform on AWS, supporting training, fine-tuning, inference, RAG, and cost optimization across the organization's AI initiatives. You'll partner with applied AI and research teams to productionize prototypes, maintain multi-provider LLM reliability (OpenAI, Google, Anthropic), and operate AI safety guardrails and authorization layers.… The role requires production-scale ML infrastructure experience on AWS (ECS, SageMaker, GPU fleets), deep knowledge of model serving and inference optimization, hands-on language model training, and demonstrated cost discipline across AI workloads.

San DiegoLast seen 19 days ago
$116,800 – $175,200 · Posted 20 days ago

Design, build, and operate enterprise-scale Databricks lakehouse platforms for a global semiconductor company. Lead AI-native development practices, machine learning enablement, and AIOps automation while mentoring engineering teams and establishing technical standards.… Drive platform modernization, cloud-native practices, and intelligent data/AI infrastructure at scale. Requires 4–5 years of hands-on Databricks experience, strong Python/SQL/Spark skills, MLOps expertise, and proven technical leadership.

San DiegoLast seen 18 days ago
Posted 20 days ago

Design and review data engineering solution architectures on Databricks and AWS, aligning technical decisions with business goals. Lead enterprise-scale data migration and modernization programs, providing architectural guidance and technical leadership across client and internal teams.… Demonstrate hands-on expertise in PySpark, SQL, data warehousing, CI/CD pipelines, and data governance frameworks. Mentor team members and foster a knowledge-sharing culture while driving smooth project execution and transition.

San DiegoLast seen 18 days ago
Posted 20 days ago

The role seeks a Databricks & AWS Data Engineering Architect to lead enterprise-scale data migration and modernization programs for a banking client. The architect will design and implement data solutions using Databricks, PySpark, SQL, and AWS, with responsibility for CI/CD pipelines, data governance frameworks, and data quality monitoring.… Required skills include hands-on expertise in data engineering, data warehousing, and data architecture, along with leadership capabilities to manage stakeholders and drive engineering best practices across the organization.

San DiegoLast seen 18 days ago
Posted 20 days ago

Site Reliability Engineer responsible for monitoring system health, detecting issues before they escalate, and owning incident response and debugging end-to-end. You'll build logging and observability tooling, automate deployment pipelines, manage capacity planning, and drive platform reliability at scale.… The role requires 5+ years in SRE/DevOps, hands-on production incident response, and strong communication during live troubleshooting. You'll work with AWS, Terraform, Datadog, Python/Bash scripting, and modern infrastructure-as-code practices.

San DiegoLast seen 18 days ago
$125,000 – $150,000 · Posted 21 days ago

Software Developer III will design, develop, test, integrate, and deploy software for the Application Arsenal platform, translating user stories and operational requirements into secure, maintainable solutions. The role requires hands-on coding in languages like Python, Java, C#, Go, or C++; implementing modern architectural patterns (microservices, event-driven, serverless); and managing the complete software delivery lifecycle including CI/CD pipelines, automated security testing (SAST, DAST, SCA), and containerization.… The developer will support Kubernetes and cloud deployments, troubleshoot issues across laboratory and operational environments, and collaborate with systems engineers, cybersecurity teams, and government stakeholders. An active Top Secret or TS/SCI clearance with Tier 5 investigation, Security+ certification, and minimum 5 years of Agile software development experience are required.

San DiegoLast seen 19 days ago
$174,000 – $185,000 · Posted 21 days ago

This is a senior full-stack engineer role responsible for architecting and building internal platforms and cloud infrastructure that serve operations, finance, lab, and external customers. The engineer will design data models and APIs connecting lab systems, business software (NetSuite, ATS, CRM), and internal dashboards; build robust cloud infrastructure managing provisioning, cost, reliability, and security; and set technical direction on architecture, code quality, CI/CD, and observability across a growing engineering org.… The role requires 10+ years of full-stack production experience, deep hands-on cloud infrastructure knowledge, modern frontend (React) and backend skills, and ability to work directly with non-engineering stakeholders to translate ambiguous business requirements into shipped software. Over time, the engineer will build and mentor a team.

San DiegoLast seen 19 days ago
$116,800 – $175,200 · Posted 21 days ago

Lead the design, build, and operation of enterprise-scale Databricks Lakehouse platforms at Qualcomm, driving AI-native development, MLOps practices, and cloud modernization. This hands-on Staff IT Engineer role requires 4–5 years of Databricks experience and 8–12 years total IT background, with responsibility for platform architecture (Delta Lake, Unity Catalog, Workflows), machine learning enablement (MLflow, Mosaic AI), AIOps automation, and technical mentoring of engineers.… You will establish engineering standards, optimize platform performance and cost, implement observability and self-healing automation, and drive adoption of generative AI tools and intelligent automation across development teams.

San DiegoLast seen 20 days ago
$207,000 – $300,000 · Posted 21 days ago

As a Forward Deployed Engineer IV in Applied AI at Google Cloud, you will lead end-to-end engineering for production conversational AI systems, transforming prototypes into scalable agentic workflows deployed at customer sites. You'll architect multi-agent systems using frameworks like ReAct, optimize RAG implementations, debug agent logic across microservices, and connect AI systems to enterprise infrastructure.… The role requires 8+ years of software development (Python or similar), experience with infrastructure-as-code (Terraform), cloud platform deployment (GCP), and hands-on full-stack application development. You'll travel up to 50% and work directly with customers to establish production AI journeys while identifying patterns that feed back into product improvements.

San DiegoLast seen 19 days ago
$155,380 – $221,416 · Posted 21 days ago

Director-level role managing infrastructure and operations for a healthcare systems company. Responsible for leading on-premises data center and enterprise facilities infrastructure, overseeing Windows/Linux servers, virtualization (VMware), SAN/NAS storage, disaster recovery, cloud platforms (Azure, AWS, GCP), networking (LAN/WAN, SD-WAN, firewalls, DNS, DHCP), identity and access management, and infrastructure monitoring.… Requires 10 years of subject-matter expertise, 8 years supervisory experience, and deep knowledge of IT service management, infrastructure automation, and compliance.

San DiegoLast seen 19 days ago