← Back to results

monitoring jobs in San Diego

$105,400 – $207,800 · Posted today

As L3 production support lead for ClaimBeacon, a Guidewire-integrated agentic AI platform, you'll own end-to-end troubleshooting and resolution of production bugs, data pipeline failures, AI agent issues, and small enhancements across the full stack (Java, Python, Flutter). You'll reproduce issues using logs and monitoring tools, decide what escalates to core engineering, and ensure fixes maintain data/AI governance and compliance standards.… The role requires 5+ years of software engineering or production support experience with proven end-to-end ownership of production issues, plus comfort working across frontend, backend, data, and infrastructure in a live, high-pressure environment.

San DiegoLast seen today
$90,000 – $107,000 · Posted 1 day ago

The Solution Architect I designs and documents cloud-native solutions using Azure and .NET, bridging business requirements with technical execution. This early-career role partners with engineering, product, and DevOps teams to translate needs into scalable, secure architectures while guiding implementation.… Responsibilities include creating architecture artifacts (diagrams, ADRs, specifications), applying modern patterns (API-first, event-driven, agentic AI), and ensuring solutions meet reliability, observability, security, and cost goals. Required: 3–5 years in software development or systems engineering, hands-on Azure and .NET experience, understanding of DevOps/CI/CD and observability principles.

San DiegoLast seen today
$79,040 – $120,640 · Posted 4 days ago

Design, build, and operate secure, scalable DevOps platforms at enterprise scale, with deep expertise in Kubernetes, Jenkins, Docker, and CI/CD infrastructure. Manage containerized workloads, implement monitoring and observability, provision and operate AWS infrastructure, and ensure platform reliability through incident response and capacity planning.… The role requires 3+ years of IT experience (or 5+ without a degree) with strong Linux administration, hands-on Kubernetes and Docker expertise, Jenkins infrastructure administration at scale, and CI/CD platform knowledge. This is a hands-on engineering role focused on platform reliability, security hardening, and developer enablement.

San DiegoLast seen 2 days ago
$189,000 – $246,000 · Posted 4 days ago

Director-level hands-on technical leader responsible for designing, implementing, and operating enterprise infrastructure across AWS, Azure, and Microsoft 365 environments. Will personally architect and troubleshoot complex hybrid cloud, networking, identity, endpoint, and workplace solutions while leading infrastructure teams, managing vendors, budgets, and roadmaps.… Must establish engineering standards for Zero Trust, disaster recovery, and business continuity while partnering across security, application, data, and business teams. Requires 15+ years of progressive infrastructure and platform engineering experience, 7+ years leading technical teams, and deep expertise in AWS services, Microsoft 365 (Entra ID, Exchange, Teams, SharePoint, Intune, Defender), Azure, Infrastructure as Code, and automation.

San DiegoLast seen 2 days ago
Posted 6 days ago

Senior administrator responsible for day-to-day support, maintenance, and configuration of SAP GRC (Governance, Risk, and Compliance) systems, including Access Controls and Process Controls. Design, develop, test, and implement SAP security roles while ensuring compliance with segregation-of-duties requirements and conducting system audits.… Troubleshoot functional and technical issues, perform performance tuning, manage HANA user security, and configure Single Sign-On across ABAP and Java stacks in S/4HANA hybrid environments.

San DiegoLast seen 4 days ago
Posted 7 days ago

This role requires an experienced architect to design and implement large-scale integrations between Siemens Teamcenter PLM and enterprise applications (ERP, MES, CAD/ECAD, data warehouse, supplier portals). The architect will own integration strategy, define interface specifications and patterns (real-time, batch, event-driven, API-based), review code for quality and scalability, and coordinate across middleware teams and multi-vendor environments during transformation programs.… Core competencies include Teamcenter integration frameworks (T4EA, T4S, T4O, ITK, SOA, REST/SOAP), ERP systems (Oracle, SAP), CAD integrations (Creo, Cadence), and ability to translate business requirements into scalable technical designs.

San DiegoLast seen 5 days ago
$190,000 – $280,000 · Posted 11 days ago

Senior Staff Lead Site Reliability Engineer to establish and mature SRE practices across Hivemind's cloud infrastructure and platform services. You will define reliability targets (SLIs/SLOs), build observability systems, lead incident response and root-cause analysis, mentor teams on reliability-first practices, and develop automation to reduce manual operational work.… The role requires 7+ years in SRE or infrastructure engineering, hands-on experience operating production services in AWS or equivalent cloud environments, expertise with containerized/distributed systems, infrastructure-as-code, and operational tooling development in Python or Go.

San DiegoLast seen 9 days ago
$143,000 – $215,000 · Posted 12 days ago

Senior Platform Engineer to design, build, and operate a cloud-native compute platform on AWS and Kubernetes. You will own platform security, scalability, and evolution through strategic initiatives and migration efforts, lead incident response and troubleshooting, and provide technical leadership and mentorship.… Requires 8+ years in platform/SRE/infrastructure engineering with deep AWS expertise, production Kubernetes experience, hands-on CI/CD (GitHub, GitHub Actions, Argo CD), Infrastructure as Code (Terraform/CDK), Kubernetes networking knowledge, and production observability/monitoring skills.

San DiegoLast seen 10 days ago
$175,000 – $195,000 · Posted 12 days ago

Senior or Staff Cloud Infrastructure Engineer responsible for designing and operating Crucible, a manufacturing-network software platform running across commercial cloud, government cloud, and on-premises environments. You will own infrastructure-as-code (Terraform, Helm, GitOps), Kubernetes production operations, cloud networking, identity and security controls, CI/CD pipelines, and observability tooling.… You'll work directly with software engineers to improve developer velocity, ensure production reliability, and support complex multi-environment deployments including GovCloud and air-gapped systems.

San DiegoLast seen 10 days ago
$90,000 – $130,000 · Posted 13 days ago

This role requires architecting and leading large-scale Teamcenter PLM integrations with enterprise systems including ERP (Oracle, SAP), MES, CAD/ECAD, data warehouses, and supplier portals. The architect will define integration strategies, design specifications, and API patterns across real-time, batch, event-driven, and middleware-based flows, while coordinating with cross-functional teams and validating implementation against approved designs.… Strong hands-on expertise in Teamcenter frameworks (T4EA, T4S, T4O, ITK, SOA, REST, SOAP, PLM XML) and coexistence architectures is essential, along with experience in ERP and CAD system integrations.

San DiegoLast seen 12 days ago
$120,001 – $160,000 · Posted 14 days ago

A Software Integration Engineer at SAIC will design, develop, and deploy a hybrid container/virtualization platform supporting enterprise-scale workloads across on-prem and cloud environments. The role encompasses infrastructure-as-code automation, Kubernetes cluster configuration, DevSecOps practices, and integration of core platform services (identity/SSO, PKI, DNS, logging/SIEM, monitoring, secrets management).… Responsibilities include automating verification/validation, establishing CI/CD pipelines and standard operating procedures, supporting GPU-enabled AI/ML workloads, and mentoring junior engineers on cloud-native best practices. The position requires 7–9+ years of experience with cloud platforms (AWS, Azure, GCP), containerization (Kubernetes, Docker), virtualization (VMware, Hyper-V, KVM), infrastructure-as-code tools (Terraform, Ansible, Puppet), and DevSecOps toolchains (ArgoCD, GitLab, Jenkins).

San DiegoLast seen 13 days ago
$149,800 – $262,200 · Posted 15 days ago

Staff SRE for ServiceNow's Government Community Cloud, providing 24/7 production support across a 3-shift team. The role combines software development, systems engineering, and networking to maintain reliability, scalability, and performance of federal infrastructure, with emphasis on automation, incident reduction, and MTTR optimization.… Requires 8+ years of related experience (or equivalent education trade-off), deep Linux knowledge, 2+ years DevOps/CI-CD and cloud experience, coding proficiency in Python/JavaScript/Ruby, database administration, and observability expertise at scale.

San DiegoLast seen 13 days ago
$90,300 – $159,900 · Posted 15 days ago

The AI Engineer designs, builds, and operates production-grade AI solutions in Azure Government (GCCH) environments, with a focus on LLM and RAG implementations. Responsibilities include testing and optimizing LLMs and model variants, benchmarking RAG systems, building MCP server integrations to connect AI agents to enterprise data, implementing agentic AI workflows using platforms like N8N and Copilot Studio, and developing modular AI microservices.… The role also involves supporting security practices (RBAC, Key Vault, data protection), monitoring and incident response for AI workloads, and helping BI and DevOps teams adopt AI capabilities toward production. A Bachelor's degree in Computer Science or related field and 1–2 years' experience with LLMs, RAG, or AI solution development are required.

San DiegoLast seen 13 days ago
Posted 15 days ago

As a Staff Site Reliability Engineer at Altium (Renesas), you will ensure reliability, availability, and performance of large-scale SaaS cloud platforms through a combination of software engineering and systems administration. You will pioneer improvements in observability (logging, monitoring, APM), develop reliability frameworks, contribute to incident response and management, and drive automation and Infrastructure as Code initiatives across multiple regions.… The role requires 6+ years of SRE/DevOps experience in large-scale environments, 3+ years of software development (ideally .NET), strong knowledge of Kubernetes, AWS, microservices, and HA architecture, plus hands-on expertise with observability tools, CI/CD platforms, and IaC tools.

San DiegoLast seen 13 days ago
$90,300 – $159,900 · Posted 15 days ago

The AI Engineer designs, builds, and operates production-grade AI solutions in Azure Government environments, focusing on LLM and RAG implementations, agentic AI workflows, and secure enterprise integrations. Responsibilities include testing and optimizing language models, benchmarking RAG systems, building MCP server integrations, implementing multi-step AI agents using platforms like N8N and Copilot Studio, and developing modular AI microservices.… The role requires hands-on support for security (RBAC, Key Vault, data protection), monitoring, and incident response in Azure Gov, with a focus on moving proofs of concept toward production deployment.

San DiegoLast seen 13 days ago
Posted 19 days ago

Shield AI seeks an experienced SRE Lead to establish and mature reliability practices across Hivemind's cloud infrastructure and platform services. This hands-on technical role involves defining SLIs/SLOs, building monitoring and alerting systems, leading incident response, and driving root-cause analysis.… The SRE Lead will mentor teams, develop operational tooling in Python or Go, and partner with product and cloud engineering teams to embed reliability-first practices into system design and infrastructure provisioning.

San DiegoLast seen 17 days ago
Posted 19 days ago

Secure the application development lifecycle and software supply chain as a Security Engineer, embedding with engineering teams to design and implement threat modeling, secure code review, SAST/DAST/SCA integration into CI/CD pipelines, and secrets management. Own dependency and supply-chain security including SCA, SBOMs, and artifact signing, while designing hardened infrastructure patterns for self-hosted applications across AWS, Azure, and on-premises.… Build durable tooling through scripting and Infrastructure-as-Code to scale security across the organization, making secure deployment and controls the default rather than exceptions.

San DiegoLast seen 17 days ago
$124,800 – $156,000 · Posted 19 days ago

As a senior technical authority on a Managed Services team, you will own the full lifecycle of cloud contact center solutions for enterprise customers using Amazon Connect. You'll build and deploy production code through automated CI/CD pipelines and Infrastructure as Code (Terraform or CloudFormation), leveraging modern Connect capabilities like Lex, Contact Lens, and Agent Assist.… You'll establish development standards, SLA targets, and quality metrics while leading 24/7 proactive monitoring, complex incident response, and root cause analysis as the senior escalation point for the team.

San DiegoLast seen 17 days ago
Posted 20 days ago

Design and review data engineering solution architectures on Databricks and AWS, aligning technical decisions with business goals. Lead enterprise-scale data migration and modernization programs, providing architectural guidance and technical leadership across client and internal teams.… Demonstrate hands-on expertise in PySpark, SQL, data warehousing, CI/CD pipelines, and data governance frameworks. Mentor team members and foster a knowledge-sharing culture while driving smooth project execution and transition.

San DiegoLast seen 18 days ago
Posted 20 days ago

The role seeks a Databricks & AWS Data Engineering Architect to lead enterprise-scale data migration and modernization programs for a banking client. The architect will design and implement data solutions using Databricks, PySpark, SQL, and AWS, with responsibility for CI/CD pipelines, data governance frameworks, and data quality monitoring.… Required skills include hands-on expertise in data engineering, data warehousing, and data architecture, along with leadership capabilities to manage stakeholders and drive engineering best practices across the organization.

San DiegoLast seen 18 days ago
Posted 20 days ago

Site Reliability Engineer responsible for monitoring system health, detecting issues before they escalate, and owning incident response and debugging end-to-end. You'll build logging and observability tooling, automate deployment pipelines, manage capacity planning, and drive platform reliability at scale.… The role requires 5+ years in SRE/DevOps, hands-on production incident response, and strong communication during live troubleshooting. You'll work with AWS, Terraform, Datadog, Python/Bash scripting, and modern infrastructure-as-code practices.

San DiegoLast seen 18 days ago
Posted 22 days ago

BAE Systems seeks a Platform DevSecOps Engineer to lead infrastructure and deployment initiatives for a defense/government software team. The role involves designing and maintaining on-premises infrastructure, developing CI/CD pipelines, automating provisioning and configuration management, implementing security measures, and troubleshooting production incidents.… The candidate must have proven DevSecOps experience, hands-on expertise with containerization (Kubernetes, Docker), automation tools (Ansible, Puppet, Chef), scripting (Python, Bash, PowerShell), CI/CD platforms (GitLab), monitoring frameworks, and Linux system administration, while working in a hybrid environment within a fast-paced, dynamic team.

San DiegoLast seen 21 days ago
$174,000 – $205,000 · Posted 24 days ago

The Senior DevOps Engineer will design, build, and maintain cloud infrastructure and CI/CD automation across Azure and AWS, managing containerized microservices at scale with Zero-Trust security patterns. The role requires architecting infrastructure-as-code solutions, deployment pipelines, and platform tooling to ensure high availability and reliability for cloud-native SaaS applications.… This position bridges development and operations, collaborating with engineering teams to implement automation and research innovative solutions. You'll need 5+ years of software engineering and DevOps experience with demonstrable expertise in cloud platforms, Kubernetes orchestration, CI/CD tools, and container technologies.

San DiegoLast seen 22 days ago
$174,000 – $205,000 · Posted 24 days ago

A Senior DevOps Engineer will design, build, and maintain cloud infrastructure and CI/CD pipelines for a clinical-trial SaaS platform on Azure and AWS. The role spans infrastructure-as-code (Terraform), container orchestration (Kubernetes, Docker), CI/CD automation (GitHub Actions, Jenkins, Azure DevOps), and monitoring/observability, with responsibility for zero-trust security, microservice scaling, and platform reliability.… The engineer will collaborate with software teams to architect cloud-native deployments and research emerging technologies to enhance platform performance. The position also involves some server-side application development support using Node.js and Python, and familiarity with LLMs for DevOps automation is a plus.

San DiegoLast seen 23 days ago
$123,000 – $185,000 · Posted 24 days ago

Staff Platform Engineer at Shield AI will design, build, and operate highly available infrastructure platforms supporting mission-critical systems across cloud and on-prem environments. Responsibilities include improving CI/CD pipelines, developer platform tooling, and automation frameworks; implementing scalable solutions for compute, networking, identity, and storage; and partnering with security teams to embed controls and compliance into platform architectures.… The role requires 8+ years of hands-on platform engineering, DevOps, or SRE experience with deep expertise in Infrastructure as Code (Terraform, Ansible), Linux systems, distributed systems, and scripting/programming languages (Python, Go, Bash, PowerShell).

San DiegoLast seen 22 days ago