← Back to results

grafana jobs in San Diego

$120,001 – $160,000 · Posted 1 month ago

Design, develop, and deploy a hybrid container/virtualization platform serving enterprise workloads across on-prem and cloud environments. Integrate core platform services (identity, PKI, DNS, logging, monitoring, secrets management) and support GPU-enabled AI/ML workloads.… Establish CI/CD pipelines, Infrastructure-as-Code automation, and standard operating procedures for configuration management, verification, and patch management. Lead cross-functional teams on containerization, DevSecOps, incident response, and ATO compliance efforts.

San DiegoLast seen 1 month ago
$120,001 – $160,000 · Posted 1 month ago

Design, develop, and deploy a hybrid container/virtualization platform supporting enterprise-scale workloads across on-prem and cloud environments. Own infrastructure automation, CI/CD integration, monitoring, security compliance, and platform service integration (identity, PKI, logging, secrets management).… Collaborate with cross-functional teams on incident response, vulnerability remediation, and mentorship of junior engineers in cloud-native and containerization best practices.

San DiegoLast seen 1 month ago
$100,000 – $170,000 · Posted 1 month ago

A Platform Engineer role focused on deploying, managing, and optimizing critical infrastructure for defense and government clients. Responsibilities include implementing system automation with Ansible and Puppet across RHEL and Windows platforms, deploying and maintaining Kubernetes clusters with RKE2 and Rancher, and managing open-source applications (NiFi, Kafka, Zookeeper, MinIO, Keycloak, Longhorn, Grafana, Prometheus, Loki, Promtail, GitLab).… The role requires 7+ years of hands-on enterprise Linux and Windows systems management, strong troubleshooting skills in distributed environments, and an active Top Secret/SCI clearance.

San DiegoLast seen 1 month ago
$200,001 – $240,000 · Posted 1 month ago

Design, build, and maintain real-time data ingestion pipelines that reliably stream data from diverse sources into a production data platform. You will ensure data quality, observability, and scalability while monitoring pipeline health, building resilient systems with proper error handling and backpressure strategies, and automating monitoring and alerting for timeliness and data issues.… Partner with data engineers, platform engineers, and analytics teams to configure pipelines for reliability, perform root cause analysis on outages, and work with security teams on encryption and data governance. Proficiency required in Python and bash scripting, plus hands-on experience with tools like Kafka, NiFi, Flink, Spark, Snowflake, Grafana, and Prometheus.

San DiegoLast seen 1 month ago
$154,000 – $231,000 · Posted 1 month ago

Design, build, and operate an enterprise Kubernetes platform at scale across Qualcomm, establishing organizational standards for cluster operations, GitOps workflows with Argo CD, and security policies using Kyverno and Cilium. Own platform reliability, observability (Datadog, Prometheus/Grafana), and cost efficiency while serving as technical authority for security governance and multi-cloud strategy.… Lead cross-organizational teams without direct authority, mentor staff engineers, drive technology evaluation across the CNCF landscape, and participate in on-call incident response including 24/7 coverage.

San DiegoLast seen 21 days ago
$200,001 – $240,000 · Posted 1 month ago

Design, build, and maintain real-time data ingestion pipelines that reliably stream data from multiple sources into a centralized data platform. You will ensure data quality, observability, and low-latency delivery while collaborating with data engineers, platform engineers, and analytics teams.… Responsibilities include building resilient pipelines with error handling, automating monitoring and alerting for data issues, performing root cause analysis on outages, and partnering with security teams on encryption and data governance. You must demonstrate proficiency in Python and bash scripting, familiarity with tools like Kafka, NiFi, Spark Streaming, Snowflake, and Elasticsearch, and 3+ years of experience in data engineering or DevOps roles supporting production pipelines.

San DiegoLast seen 1 month ago
$188,500 – $255,000 · Posted 1 month ago

Staff Software Engineer to architect and lead Intuit's next-generation One Logging system, designing high-scale, mission-critical logging pipelines for on-premise and public cloud (AWS, GCP) deployments. Own end-to-end platform evolution from log ingestion through routing, storage, and cost optimization, including components like FELS, Kinesis/CloudWatch integrations, and Log Router.… Lead modernization of legacy tooling, design edge collection agents (Fluent Bit, OIL sidecar), and drive automation initiatives using MCP Server capabilities. Provide technical leadership, mentorship, and cross-functional collaboration to align logging platform capabilities across the organization.

San DiegoLast seen 1 month ago
$92,400 – $148,800 · Posted 1 month ago

A Senior Site Reliability Engineer at SHEIN will own and operate mission-critical, large-scale distributed systems (Kubernetes, Kafka, Elasticsearch, Redis, APISIX, Nginx) running 24/7/365, participating in on-call rotations and driving fast incident response using AI-assisted log analysis and anomaly detection. The role requires strong software engineering expertise in Python or Go, deep Linux and networking knowledge, and hands-on experience with observability platforms (Prometheus, Grafana) and configuration management tools.… Responsibilities include designing resilient monitoring and alerting infrastructure, automating operational workflows to eliminate toil, capacity planning, and collaborating with global teams to improve system reliability and performance.

San DiegoLast seen 4 days ago
$150,000 – $220,000 · Posted 1 month ago

Shield AI seeks a Staff Engineer, DevOps to own and maintain the build system for autonomous aircraft (XBAT/VBAT), with deep expertise in C++ and CMake. You will improve developer velocity by reducing build/test cycle times, integrating CI pipelines with HIL and simulation environments, and implementing monitoring and dashboards for build system performance.… The role requires 7+ years of industry experience (or 5+ with advanced degree), proficiency with containers/orchestration, scripting, CI/CD tools, and cloud operations across Azure/AWS/GCP. You'll collaborate with software engineers to resolve failures, set best practices, and leverage agentic AI solutions to identify root causes and accelerate development cycles.

San DiegoLast seen 7 days ago
Posted 1 month ago

Design and own cloud infrastructure foundations, core platform services, and observability systems that enable reliable deployment across customer environments (on-premise, cloud, or hybrid). Build production-grade Kubernetes infrastructure, internal APIs, deployment pipelines, and monitoring/alerting layers using infrastructure-as-code.… Instrument service-level objectives and health signals to ensure measurable reliability and reproducible, secure deployments across all environments.

San DiegoLast seen 11 days ago
Posted 1 month ago

Build and operate the data and ML infrastructure powering an AI platform for materials science, owning both sides: data pipelines that ingest and curate large-scale scientific output into training-ready formats, and model packaging, serving, monitoring, and CI/CD systems that move models safely from research to production across customer environments. You will design data ingestion and transformation workflows, implement validation and quality gates, package and version models with reproducible builds, run models through batch and online inference with safe rollout and rollback, monitor for drift and degradation, and build observability and internal tooling for engineering and science teams.… The role requires 6+ years shipping production software with deep expertise in data systems, ML infrastructure, containers, orchestration, and observability.

San DiegoLast seen 11 days ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve infrastructure, deployment systems, and cloud environments powering Clinically AI's healthcare AI platform. Design and maintain GCP infrastructure with focus on GKE, Kubernetes, and secure multi-environment deployments; build Infrastructure-as-Code using Terraform and Helm charts; optimize CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for services, AI pipelines, and high-throughput workloads.

San DiegoLast seen 1 month ago
$125,000 – $160,000 · Posted 1 month ago

Own and evolve the cloud infrastructure, deployment systems, and CI/CD pipelines powering Clinically AI's healthcare platform on GCP. Design and maintain cloud infrastructure using GKE, Terraform, and Helm; build reliable CI/CD pipelines with GitHub Actions; implement observability, security best practices, and incident response processes.… Work closely with Backend, AI, and Product teams to support scalable infrastructure for AI pipelines, real-time processing, and multi-environment deployments. Operate with significant autonomy across cloud-native architecture while optimizing for performance, reliability, cost, and compliance.

San DiegoLast seen 1 month ago
$108,000 – $180,000 · Posted 1 month ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper).… Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 13 days ago
$120,000 – $160,000 · Posted 1 month ago

This role provides Tier 2 technical support for Windows and Linux systems, network infrastructure, and enterprise applications in a classified DoD environment. The administrator will configure and maintain operating systems, troubleshoot hardware and software issues, support distributed systems (Hadoop, Kafka, HBase/Accumulo, Spark), and utilize automation tools including Bash, PowerShell, Python, Ansible, Puppet, and Terraform.… Responsibilities include maintaining compliance with DoD and Intelligence Community security requirements, collaborating with engineering and cybersecurity teams, and documenting technical procedures. An active TS/SCI clearance and DoD 8570/8140 IAT Level I certification are required.

San DiegoLast seen 1 month ago