← Back to results

observability jobs in San Diego

$85,000 – $100,000 · Posted 1 day ago

As Full Stack Engineer at Nucleus Biologics, you will develop and maintain software tools for scientists and manufacturing teams. Responsibilities include managing AWS infrastructure (EC2, RDS), maintaining Docker containers and Nginx, handling PostgreSQL schema and migrations, developing Python backends with FastAPI/Django/Flask, building React/TypeScript frontends, and resolving customer and internal issues. You need strong production experience across the full stack: Python with modern web frameworks, React with current data-fetching patterns, PostgreSQL administration, and hands-on DevOps with Docker, Linux, and AWS.

San DiegoLast seen today
$230,000 – $311,000 · Posted 1 day ago

Lead product strategy and roadmap for Intuit's next-generation Agentic AI Platform, designing core execution frameworks, agent libraries, and SDKs that enable complex reasoning and autonomous task execution across TurboTax, QuickBooks, Mailchimp, and Credit Karma. Define extensible platform architecture, drive multi-step planning and orchestration capabilities, and ensure agent observability and trust at scale. Partner closely with engineering and architects to standardize agent development, deployment, and management while mentoring other product managers on technical fluency and platform thinking.

San DiegoLast seen today
Posted 1 day ago

Lead the design, development, and deployment of enterprise-scale Generative AI and LLM solutions across Azure and AWS environments. Define AI architecture frameworks, standards, and governance practices while architecting RAG systems, agentic AI, multi-agent systems, and vector database solutions. Partner with business leaders and engineering teams to translate business requirements into scalable, secure, and responsible AI platforms. Mentor architects and data science teams while evaluating emerging AI technologies and driving innovation initiatives.

San DiegoLast seen today
$122,600 – $177,900 · Posted 1 day ago

SHEIN is seeking a Senior Data Engineer to build and productionize GenAI/LLM solutions for data engineering workflows, including code assistance, metadata discovery, and knowledge access. The role focuses on developing RAG and agentic workflows, automating incident triage and log analysis, and integrating AI capabilities into internal developer tools and services. You'll own scoped projects from problem definition through production support, requiring 3+ years of hands-on production experience with strong Python/SQL fundamentals, proven GenAI/LLM application development, and practical knowledge of databases, APIs, data pipelines, and distributed systems.

San DiegoLast seen today
$150,000 – $185,000 · Posted 2 days ago

Senior hands-on software engineer building the core orchestration, data systems, and internal tooling for an automated chemistry platform. You'll own problems end-to-end across the full stack—Python backend, TypeScript/React frontend, databases, APIs, and distributed systems—designing services, modeling data flows, and building interfaces that connect synthesis requests from intake through execution. You'll work closely with automation engineers, chemists, and ML teams to ship production features, establish engineering practices, and become a primary owner of major platform components. The role prioritizes judgment and problem-solving over language expertise, expecting fluency across the stack and comfort reaching for languages like Rust, Go, or C++ when warranted.

San DiegoLast seen today
$150,900 – $413,600 · Posted 3 days ago

This role requires a senior technical architect with 15+ years of Integrated Eligibility System Architecture and cloud infrastructure expertise to design and govern end-to-end data center infrastructure across compute, storage, and networking. The architect will define standards and patterns, lead infrastructure modernization to AWS, and provide strategic guidance on scalable, secure, AI-enabled environments. Deep knowledge of cloud-native engineering, Infrastructure as Code, DevOps, and compliance/security requirements is essential, with particular experience in agentic AI and serverless Integrated Eligibility systems preferred.

San DiegoLast seen 1 day ago
$202,500 – $274,000 · Posted 3 days ago

Staff Data Engineer role at Credit Karma building petabyte-scale data infrastructure and streaming platforms on GCP. You will design and operate Kafka-based streaming infrastructure, build cloud-native data pipelines using Dataflow/Beam, Flink, and Spark, and create persistence frameworks across Spanner, MySQL, and BigQuery. The role requires 7+ years backend/data systems experience in JVM languages (Scala/Java), expertise with high-throughput distributed systems, and deep knowledge of streaming platforms, Apache Beam, CDC patterns, encryption, and data governance frameworks. You'll also integrate AI/ML and generative AI technologies including LLMs, RAG, semantic search, and knowledge graphs into the data platform.

San DiegoLast seen 1 day ago
$122,600 – $177,900 · Posted 3 days ago

Senior Data Engineer role focused on building and productionizing GenAI/LLM solutions for data engineering workflows. You will develop RAG and agentic workflows, automate incident triage and root-cause analysis, integrate AI capabilities into internal developer tools, and own projects from problem definition through production support. Required: 3+ years of production systems experience, strong Python/SQL, hands-on GenAI/LLM application building, and software engineering fundamentals including testing, CI/CD, and monitoring.

San DiegoLast seen 1 day ago
$264,000 – $330,000 · Posted 4 days ago

Principal Machine Learning Engineer at AppFolio to architect and lead mission-critical AI systems across the Realm-X property management platform. You'll design advanced AI agentic systems combining reasoning, planning, and execution; establish ML platform primitives for end-to-end production workflows; and drive the transition toward autonomous property management using LLMs, fine-tuning, and reinforcement learning. Requires 10+ years building software systems with deep expertise in traditional ML, deep learning, generative AI, LLM post-training (SFT, RLHF, DPO, RL), and production ML at scale—plus a Master's or Ph.D. in Computer Science or related field.

San DiegoLast seen 2 days ago
$92,400 – $148,800 · Posted 4 days ago

A Senior Site Reliability Engineer at SHEIN will own and operate mission-critical, large-scale distributed systems (Kubernetes, Kafka, Elasticsearch, Redis, APISIX, Nginx) running 24/7/365, participating in on-call rotations and driving fast incident response using AI-assisted log analysis and anomaly detection. The role requires strong software engineering expertise in Python or Go, deep Linux and networking knowledge, and hands-on experience with observability platforms (Prometheus, Grafana) and configuration management tools. Responsibilities include designing resilient monitoring and alerting infrastructure, automating operational workflows to eliminate toil, capacity planning, and collaborating with global teams to improve system reliability and performance.

San DiegoLast seen 2 days ago
$95,500 – $95,500 · Posted 4 days ago

Design and build production AI solutions including copilots, agents, RAG applications, and workflow automations on Azure OpenAI, Claude, and Microsoft platforms. Write custom code in Python, JavaScript/TypeScript, or C# and leverage low-code platforms like Power Automate and Logic Apps where appropriate, following solid engineering practices (version control, testing, CI/CD, AI output evaluation). Integrate solutions with enterprise systems and APIs with security-first design, own projects end-to-end from deployment through adoption and improvement, and work with security, architecture, governance, and compliance teams to ensure production readiness. Communicate clearly to both technical and business audiences, document designs, and coach teammates on AI capabilities, limits, and risks.

San DiegoLast seen 3 days ago
$128,100 – $192,100 · Posted 5 days ago

Staff-level data engineer who designs, develops, and maintains scalable ETL/ELT pipelines using Databricks, PySpark, and Python for enterprise analytics and AI use cases. Builds curated data layers following Lakehouse Medallion architecture, develops Databricks-native applications (notebooks, dashboards, APIs), and optimizes Apache Spark workloads for performance and cost efficiency. Owns production data pipelines end-to-end, ensures data quality and reliability, and acts as a technical leader mentoring teams on data engineering best practices and modern Databricks capabilities.

San DiegoLast seen 4 days ago
$180,000 – $230,000 · Posted 5 days ago

Zensar is seeking an experienced AI Architect to design, develop, and deploy enterprise-scale AI and Generative AI solutions. The role requires deep expertise in Large Language Models (Claude, Gemini, OpenAI), multi-cloud platforms (Azure, AWS), and enterprise architecture, with 10+ years in software/cloud architecture and 5+ years designing AI/ML solutions. Responsibilities include defining AI strategy and architecture standards, designing end-to-end GenAI systems (RAG, agentic AI, multi-agent systems), architecting secure infrastructure across Azure and AWS, establishing governance and responsible AI frameworks, and leading adoption programs. The ideal candidate will mentor teams, evaluate emerging technologies, and translate business requirements into scalable, production-grade AI systems.

San DiegoLast seen 4 days ago
Posted 5 days ago

Senior/Staff Data Engineer responsible for building production data pipelines, APIs, and BI tools that serve Agoda's contact center operations. Will design and maintain dashboards, metrics, and analytics platforms using SQL, Python, Java, Golang, or TypeScript; work with data warehouses like BigQuery, Snowflake, or Redshift; and orchestrate ETL/ELT workflows using tools such as Airflow or dbt. At Staff level, defines roadmap for key data platforms, mentors junior engineers, and influences cross-team data architecture to scale with business growth.

San DiegoLast seen 4 days ago
$114,266 – $171,400 · Posted 6 days ago

Senior Automated Test Engineer designing and scaling advanced automated testing and CI/CD infrastructure for autonomous unmanned aircraft and robotic platforms. The role involves developing test frameworks, CI/CD pipelines, and verification tooling across distributed systems including flight code, simulation environments, ground control stations, and hardware-in-the-loop systems. You will work with Python automation, UI/API testing in Go and TypeScript, GitLab CI/CD, and collaborate with software, autonomy, and systems engineering teams to ensure mission-critical reliability and rapid iteration. Experience with autonomy, aerospace, safety-critical systems, and AI-assisted development is valued.

San DiegoLast seen 4 days ago
$141,200 – $278,300 · Posted 6 days ago

Lead the design and delivery of enterprise AI platforms and applications on Google Cloud, leveraging Vertex AI, Gemini, and cloud-native technologies. Design, fine-tune, and govern LLM solutions; build RAG and agentic systems; and define end-to-end architectures spanning data pipelines, feature engineering, model lifecycle, APIs, and MLOps/LLMOps. Architect cloud-native applications on GKE, Cloud Run, and managed services while implementing security, governance, and production-grade monitoring for AI/ML systems at scale.

San DiegoLast seen 4 days ago
$149,800 – $262,200 · Posted 7 days ago

Staff Software Engineer role designing and implementing cloud-native platform components and services, with primary focus on Kubernetes workloads, distributed systems, and infrastructure automation. Requires 5+ years hands-on Kubernetes experience, proficiency in Go or systems languages, and demonstrated expertise building operators, controllers, and platform services. Will lead well-scoped projects from design through production, contribute to reliability and scalability, mentor junior engineers, and participate in on-call operations.

San DiegoLast seen 4 days ago
$162,600 – $244,000 · Posted 7 days ago

As a Datacenter AI Systems and Solutions Engineer at Qualcomm, you will research, develop, and optimize end-to-end AI/ML solutions that integrate Qualcomm's AI inference accelerators with system software and ecosystem components. You will lead the design and deployment of production-ready Generative AI and LLM applications, perform model benchmarking and performance analysis, and serve as a technical lead for customer engagements on AI model optimization and inference tuning. The role requires deep expertise in AI systems architecture, MLOps practices, and large-scale distributed systems, with hands-on proficiency in Python, ML frameworks, containerization, and orchestration platforms. You will drive system-level requirements, hardware/software co-design, and influence product direction through performance analysis and optimization strategies.

San DiegoLast seen 4 days ago
$220,203 – $275,254 · Posted 7 days ago

Lead a team of cloud engineers building the nervous system for Brain Corp's fleet of 30,000+ autonomous mobile robots. Own the platform's reliability, roadmap, and architecture while managing team hiring, onboarding, and career development. Drive high-performance delivery across robot and customer-facing interfaces, balance competing priorities (performance, cost, reliability, scalability), and partner with principal and staff engineers on long-term platform strategy. Stay hands-on enough to make sound architectural decisions and earn senior engineer trust during a period of significant growth.

San DiegoLast seen 4 days ago
$220,203 – $275,254 · Posted 7 days ago

Lead a team of cloud engineers building the backend platform that powers Brain Corp's global fleet of 30,000+ autonomous mobile robots. Own the platform's reliability, roadmap, and architecture as it scales to support fleet operations and customer-facing applications. Balance hands-on technical leadership with people management—mentor engineers, drive hiring and onboarding, own production reliability and incident response, and partner with principal/staff engineers on distributed-systems architecture. Navigate tradeoffs between performance, cost, reliability, and scalability while shipping features that robots in the field depend on.

San DiegoLast seen 4 days ago
$179,200 – $268,800 · Posted 7 days ago

Qualcomm seeks an experienced Staff Product Manager to define and drive the Linux platform roadmap across Snapdragon-based AI PCs, Edge AI systems, and enterprise computing solutions. The role spans product strategy, Linux distribution enablement, upstream kernel and driver strategy, and AI developer platform support. You will gather customer and ecosystem requirements, create product roadmaps, collaborate with OEMs, ISVs, Linux distribution partners, and open-source communities to deliver optimized Linux experiences on Snapdragon platforms. Success requires translating market needs into platform requirements, managing cross-functional alignment across hardware, firmware, kernel, and AI software layers.

San DiegoLast seen 4 days ago
$286,000 – $429,000 · Posted 8 days ago

Sr. Director role overseeing SIE's data and AI platform ecosystem, including data infrastructure (streaming, ingestion, warehouse), experimentation platforms, and governance tools. Responsible for defining long-term vision, leading global engineering teams, and ensuring platforms process massive data volumes to deliver insights across gaming, content creation, and enterprise operations. Must drive scalability, cost efficiency, security, and operational excellence while collaborating with senior leadership to align platform strategy with business objectives. Requires proven Sr. Director-level experience and thought leadership in platform architecture, real-time data processing, and ML/AI integration.

San DiegoLast seen 6 days ago
$177,300 – $265,900 · Posted 8 days ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues. Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 6 days ago
$187,363 – $265,900 · Posted 8 days ago

Design and architect next-generation ML inference infrastructure for globally distributed, multi-tenant model serving with high availability, scalability, and cost efficiency. Lead the development of low-latency, high-throughput inference systems supporting computer vision and multimodal models (CNNs, segmentation, object detection) using Scala, Java, and Go. Build large-scale distributed systems with reactive frameworks, integrate enterprise feature stores, and extend CI/CD pipelines (GitHub Prow, Pulumi) with automation and policy enforcement. Design advanced observability frameworks, optimize ML algorithms for performance, mentor engineers, and ensure MLOps and compliance standards (GDPR, SOC2) across the platform.

San DiegoLast seen 6 days ago
Posted 8 days ago

The AI Architect will lead the design, development, and deployment of enterprise-scale generative AI solutions across Azure and AWS, partnering with business and engineering leaders to define AI strategy and architecture standards. The role requires deep expertise in modern LLMs (Claude, Gemini, OpenAI), multi-cloud AI platforms, and enterprise architecture, with 10+ years in software/cloud architecture and 5+ years architecting AI/ML solutions. Key responsibilities include designing end-to-end GenAI solutions leveraging RAG, agentic AI, and multi-agent systems; establishing cloud-native deployment patterns; and defining governance, compliance, and responsible AI frameworks. The ideal candidate will serve as a trusted technical advisor to executives, mentor engineering teams, and drive AI adoption through reusable frameworks and reference architectures.

San DiegoLast seen 6 days ago