← Back to results

prometheus jobs in San Diego

$150,000 – $200,000 · Posted 4 days ago

Cloud Engineer III will deploy, configure, administer, and troubleshoot cloud-based compute, storage, networking, and application infrastructure on AWS and Kubernetes. The role requires hands-on expertise in Linux systems administration, containerized applications (Docker/Kubernetes), cloud networking, IAM, security hardening, and observability platforms.… Responsibilities include supporting production Kubernetes clusters (EKS), implementing Infrastructure as Code (Terraform/CloudFormation), integrating CI/CD pipelines, automating administration with Python/Bash, and troubleshooting complex issues across cloud, container, and application layers. The position demands strong security practices, compliance knowledge for government environments, and collaboration with cross-functional technical and operational teams.

San DiegoLast seen 2 days ago
$90,000 – $125,000 · Posted 6 days ago

Full-Stack Software Engineer responsible for designing and building an on-premises AI server project focused on privacy for defense applications, with end-to-end ownership from planning through development, testing, deployment, and production support. Must be proficient in one backend language (Go, Rust, Python, C#/.NET, or Node.js), frontend frameworks (React or Vue), and Linux/Docker, with hands-on experience managing GPU workloads, LLM and image generation models, and networking/security.… This is a 100% onsite, full-time role in San Diego with significant technical decision-making authority over system architecture and quality.

San DiegoLast seen 4 days ago
Posted 6 days ago

This is a hands-on DevSecOps engineering role focused on designing and maintaining secure CI/CD pipelines, software supply-chain controls, and Kubernetes deployment standards. The engineer will work across Azure DevOps, Nexus, SonarQube, and container orchestration to integrate security and quality checks into build and release workflows, support platform engineering teams, and enable developers with documentation and automation.… Responsibilities include implementing SAST, SCA, dependency scanning, secrets management, artifact controls, and compliance alignments (NIST, CMMC, RMF, Zero Trust); troubleshooting pipeline and deployment failures; and creating runbooks and guardrails. The role requires 5+ years in DevOps/DevSecOps, hands-on Azure DevOps and Kubernetes experience, scripting proficiency, and an active Secret Clearance with eligibility for TS/SCI.

San DiegoLast seen 4 days ago
$90,000 – $125,000 · Posted 6 days ago

Design and build an on-premises AI server system focused on privacy for defense applications, with end-to-end ownership from planning through deployment and production support. You will make all technical decisions and own the system's quality and performance.… Required skills include a backend language (Go, Rust, Python, C#/.NET, or Node.js), frontend experience (React or Vue with HTML/CSS/JavaScript), Linux/Docker expertise, and AI/GPU knowledge including LLM and image generation models. U.S. citizenship and ability to obtain a DoD security clearance are mandatory.

San DiegoLast seen 4 days ago
Posted 8 days ago

Design, build, and maintain real-time data ingestion pipelines that reliably flow streaming data from multiple sources into a data platform while ensuring quality, observability, and scalability. Monitor data health, develop resilient pipelines with proper error-handling and backpressure strategies, and automate monitoring of real-time feeds to alert on timeliness, volume, and distribution issues.… Collaborate with data, platform, and software engineers to optimize pipeline configuration, perform root-cause analysis on data outages, and work with security teams on encryption and data classification. Requires 3+ years in data engineering, data operations, or DevOps roles; proficiency with Python and bash; and hands-on experience with tools like Kafka, NiFi, Spark Streaming, Snowflake, Elasticsearch, Grafana, and Prometheus.

San DiegoLast seen 6 days ago
Posted 11 days ago

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines.… The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.

San DiegoLast seen 9 days ago
Posted 11 days ago

The DevOps Engineer owns the design, deployment, and operation of production infrastructure at scale across AWS and Kubernetes, with responsibility for CI/CD pipelines, observability, and incident response. The role requires hands-on expertise in AWS services (EC2, EKS, VPC, IAM, S3, RDS), Kubernetes cluster operations, Terraform-based infrastructure-as-code, and monitoring tools like Prometheus and Grafana.… You will automate operational workflows with Python, Go, or Bash, lead incident investigations, and partner with software engineers and SREs to improve release velocity and system reliability. 3–8 years of DevOps, SRE, or platform engineering experience is expected.

San DiegoLast seen 9 days ago
Posted 12 days ago

Senior Network Engineer to design, build, and operate critical network infrastructure for cloud, AI/GPU, and high-performance computing environments. Responsibilities include managing multi-homed internet edge networks with BGP/eBGP, designing high-availability architectures with redundant firewalls and load balancing, deploying InfiniBand/GPU networks, automating configuration with Python/Ansible, and leading incident response.… Requires 7+ years of network engineering experience, strong BGP and TCP/IP expertise, hands-on work with Cisco/Arista/Fortinet platforms, and network automation skills.

San DiegoLast seen 10 days ago
$160,000 – $290,000 · Posted 14 days ago

Staff Engineer who will design, build, and operate the Forge Platform—a distributed-systems foundation serving autonomy, ML Ops, simulation, and application teams. You will own architecture and technical standards for workflow orchestration, asynchronous processing, event-driven systems, and long-running service workflows while remaining hands-on in implementation, production troubleshooting, and reliability improvement.… Required: strong production experience with distributed systems, cloud-native platforms, or backend infrastructure; fluency in Go and Python; deep understanding of failure handling, consistency, fault tolerance, and state management; and ability to turn recurring infrastructure needs into reusable platform capabilities. You will work across Kubernetes, service networking, observability, and data pipelines to enable downstream teams to move faster on mission-critical systems.

San DiegoLast seen 12 days ago
$160,000 – $290,000 · Posted 14 days ago

Staff Engineer responsible for designing, building, and operating the Forge Platform—a distributed-systems foundation serving autonomy, ML, simulation, and test teams. You will architect cloud-native backend infrastructure, establish reusable platform capabilities, define technical standards, and remain hands-on in production troubleshooting, performance analysis, and reliability improvement.… Required: deep expertise in production distributed systems, Go and Python, workflow orchestration, and the ability to translate complex architecture into usable interfaces across multiple downstream teams.

San DiegoLast seen 13 days ago
$120,001 – $160,000 · Posted 14 days ago

A Software Integration Engineer at SAIC will design, develop, and deploy a hybrid container/virtualization platform supporting enterprise-scale workloads across on-prem and cloud environments. The role encompasses infrastructure-as-code automation, Kubernetes cluster configuration, DevSecOps practices, and integration of core platform services (identity/SSO, PKI, DNS, logging/SIEM, monitoring, secrets management).… Responsibilities include automating verification/validation, establishing CI/CD pipelines and standard operating procedures, supporting GPU-enabled AI/ML workloads, and mentoring junior engineers on cloud-native best practices. The position requires 7–9+ years of experience with cloud platforms (AWS, Azure, GCP), containerization (Kubernetes, Docker), virtualization (VMware, Hyper-V, KVM), infrastructure-as-code tools (Terraform, Ansible, Puppet), and DevSecOps toolchains (ArgoCD, GitLab, Jenkins).

San DiegoLast seen 13 days ago
Posted 22 days ago

BAE Systems seeks a Platform DevSecOps Engineer to lead infrastructure and deployment initiatives for a defense/government software team. The role involves designing and maintaining on-premises infrastructure, developing CI/CD pipelines, automating provisioning and configuration management, implementing security measures, and troubleshooting production incidents.… The candidate must have proven DevSecOps experience, hands-on expertise with containerization (Kubernetes, Docker), automation tools (Ansible, Puppet, Chef), scripting (Python, Bash, PowerShell), CI/CD platforms (GitLab), monitoring frameworks, and Linux system administration, while working in a hybrid environment within a fast-paced, dynamic team.

San DiegoLast seen 21 days ago
$149,800 – $262,200 · Posted 25 days ago

Staff Software Engineer on the AI Experience Framework (AIUX) team, responsible for full-stack architectural decisions in a JavaScript/component-driven environment powering ServiceNow's AI-first user interfaces. You will own code from design through delivery, architect modular reusable component systems, set framework-level technical direction for AI integration and state management, and serve as an escalation point for production issues across multiple teams.… The role requires 8+ years of OO language experience (Java, C++, C#, Go), advanced expertise in modern UI frameworks (React, Angular, Vue, Lit), relational databases, and data structures/algorithms/design patterns at scale.

San DiegoLast seen 23 days ago
$149,800 – $262,200 · Posted 26 days ago

Staff Software Engineer role owning full-stack architecture for ServiceNow's AI Experience Framework (AIUX), building modular Lit-based web components that power conversation-first applications. Responsible for end-to-end code delivery from design through production, setting technical direction for framework-level APIs, and serving as an escalation point for production issues across multiple teams.… Requires 8+ years of OO language experience (Java, C++, C#, Go), advanced JavaScript/modern UI framework expertise (Angular, React, Vue, Lit), and deep knowledge of data structures, algorithms, design patterns, and performance optimization. Expected to mentor colleagues, coordinate cross-org architecture decisions, and architect scalable component systems for extensibility.

San DiegoLast seen 23 days ago
Posted 26 days ago

Design and build platform engineering infrastructure, tools, and automation for cloud-native environments on AWS, using Infrastructure as Code (Terraform/CDK), CI/CD, and observability practices. Develop reusable modules, APIs, CLIs, and service templates that improve developer experience and enforce security, compliance, and cost controls.… Establish monitoring, logging, alerting, SLOs, and incident response practices; participate in 24x7 on-call rotation with a focus on production reliability and automation. Requires 7+ years in DevOps, SRE, platform engineering, or cloud operations with deep hands-on experience in multi-account AWS setups, GitOps, and observability platforms like Datadog or Prometheus.

San DiegoLast seen 24 days ago
Posted 27 days ago

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production services running on AWS and Kubernetes. Responsibilities include designing highly available infrastructure with Terraform and managed AWS services, building CI/CD pipelines with GitHub Actions and Argo CD, defining SLOs and implementing observability with Prometheus and Grafana, leading incident response, and automating operational work with Python, Go, or Bash.… The role requires 3–8 years of hands-on SRE or DevOps experience, production Kubernetes expertise, strong AWS and infrastructure-as-code knowledge, and proficiency with observability and deployment strategies.

San DiegoLast seen 25 days ago
$120,001 – $160,000 · Posted 1 month ago

Software Integration Engineer responsible for designing, developing, and deploying a hybrid container/virtualization platform-as-code solution supporting enterprise-scale workloads across on-prem and cloud environments. The role involves integrating core platform services (identity, PKI, DNS, logging, monitoring, secrets management), configuring container runtimes and networking, and developing infrastructure-as-code automation.… Requires 9+ years of experience (or equivalent with advanced degree) with cloud platforms, Kubernetes, virtualization, DevSecOps practices, CI/CD tooling, and automation frameworks, plus the ability to mentor junior engineers and collaborate across security, platform, and application teams on compliance, incident response, and AI/ML workload enablement.

San DiegoLast seen 1 month ago
$233,000 – $350,000 · Posted 1 month ago

Senior Staff Engineer responsible for designing and building an MLOps platform that supports distributed AI training, reinforcement learning, and foundation model development at scale. You will architect Kubernetes-native infrastructure for GPU workloads, design self-service AI development workflows, manage the data and model lifecycle, and lead platform distribution across cloud, on-premises, and air-gapped environments.… The role requires deep expertise in modern AI frameworks (PyTorch, Hugging Face Transformers), distributed systems, GPU scheduling, and cloud-native infrastructure, with a focus on enabling researchers and engineers to move from experimentation to production rapidly.

San DiegoLast seen 1 month ago
$171,900 – $300,800 · Posted 1 month ago

Lead a team of full-stack software engineers building ServiceNow's AI Experience Framework (AIUX) — a platform delivering AI-first, conversation-first user interfaces through modular Lit-based web components. You'll manage product development, set technical direction for framework-level APIs, guide architecture decisions across the request path, and represent AIUX in cross-org discussions.… The role requires deep expertise in JavaScript and Java/C++/C#/Go, 10+ years of relevant technology experience, 5+ years managing core engineering teams, and the ability to solve complex problems spanning AI integration, performance at scale, and multi-team coordination.

San DiegoLast seen 1 month ago
Posted 1 month ago

DevOps Engineer to build, maintain, and optimize AWS cloud infrastructure and CI/CD deployment systems, working with ECS, EKS, and EC2 workloads. The role requires hands-on experience with Infrastructure as Code (Terraform), containerization (Docker/Kubernetes), monitoring tools (Grafana), and CI/CD pipelines, with growing ownership of services and systems.… You'll troubleshoot infrastructure and application issues, collaborate with engineering teams on deployment workflows, and apply AI tooling to improve infrastructure automation and operational efficiency. The position reports to a DevOps Manager and is based in San Diego on a hybrid schedule.

San DiegoLast seen 16 days ago
$132,962 – $226,035 · Posted 1 month ago

This is a DevSecOps Engineer role requiring hands-on expertise in cloud infrastructure, containerization, and CI/CD automation for a classified government project. The role demands proficiency with AWS, Kubernetes, Docker, scripting (Python, Bash, PowerShell), infrastructure-as-code tools (Ansible, Puppet, Chef), and Linux system administration.… The candidate must be able to obtain or hold a TS/SCI security clearance and will work on air-gapped and on-premises environments alongside cloud services.

San DiegoLast seen 1 month ago
$120,001 – $160,000 · Posted 1 month ago

Design, develop, and deploy a hybrid container/virtualization platform serving enterprise workloads across on-prem and cloud environments. Integrate core platform services (identity, PKI, DNS, logging, monitoring, secrets management) and support GPU-enabled AI/ML workloads.… Establish CI/CD pipelines, Infrastructure-as-Code automation, and standard operating procedures for configuration management, verification, and patch management. Lead cross-functional teams on containerization, DevSecOps, incident response, and ATO compliance efforts.

San DiegoLast seen 1 month ago
$120,001 – $160,000 · Posted 1 month ago

Design, develop, and deploy a hybrid container/virtualization platform supporting enterprise-scale workloads across on-prem and cloud environments. Own infrastructure automation, CI/CD integration, monitoring, security compliance, and platform service integration (identity, PKI, logging, secrets management).… Collaborate with cross-functional teams on incident response, vulnerability remediation, and mentorship of junior engineers in cloud-native and containerization best practices.

San DiegoLast seen 1 month ago
$100,000 – $170,000 · Posted 1 month ago

A Platform Engineer role focused on deploying, managing, and optimizing critical infrastructure for defense and government clients. Responsibilities include implementing system automation with Ansible and Puppet across RHEL and Windows platforms, deploying and maintaining Kubernetes clusters with RKE2 and Rancher, and managing open-source applications (NiFi, Kafka, Zookeeper, MinIO, Keycloak, Longhorn, Grafana, Prometheus, Loki, Promtail, GitLab).… The role requires 7+ years of hands-on enterprise Linux and Windows systems management, strong troubleshooting skills in distributed environments, and an active Top Secret/SCI clearance.

San DiegoLast seen 1 month ago
Posted 1 month ago

Lead a platform DevSecOps team at BAE Systems, designing and maintaining on-premises infrastructure with a focus on security, scalability, and reliability. You will develop CI/CD pipelines, automate infrastructure provisioning using tools like Ansible or Puppet, implement security best practices, and troubleshoot production incidents.… The role requires hands-on expertise in containerization (Kubernetes, Docker), scripting (Python, Bash, PowerShell), monitoring, and Linux administration, with the ability to bridge development and operations in a fast-paced, hybrid environment.

San DiegoLast seen 19 days ago