← Back to results

grafana jobs in San Diego

$150,000 – $200,000 · Posted 4 days ago

Cloud Engineer III will deploy, configure, administer, and troubleshoot cloud-based compute, storage, networking, and application infrastructure on AWS and Kubernetes. The role requires hands-on expertise in Linux systems administration, containerized applications (Docker/Kubernetes), cloud networking, IAM, security hardening, and observability platforms.… Responsibilities include supporting production Kubernetes clusters (EKS), implementing Infrastructure as Code (Terraform/CloudFormation), integrating CI/CD pipelines, automating administration with Python/Bash, and troubleshooting complex issues across cloud, container, and application layers. The position demands strong security practices, compliance knowledge for government environments, and collaboration with cross-functional technical and operational teams.

San DiegoLast seen 2 days ago
Posted 6 days ago

This is a hands-on DevSecOps engineering role focused on designing and maintaining secure CI/CD pipelines, software supply-chain controls, and Kubernetes deployment standards. The engineer will work across Azure DevOps, Nexus, SonarQube, and container orchestration to integrate security and quality checks into build and release workflows, support platform engineering teams, and enable developers with documentation and automation.… Responsibilities include implementing SAST, SCA, dependency scanning, secrets management, artifact controls, and compliance alignments (NIST, CMMC, RMF, Zero Trust); troubleshooting pipeline and deployment failures; and creating runbooks and guardrails. The role requires 5+ years in DevOps/DevSecOps, hands-on Azure DevOps and Kubernetes experience, scripting proficiency, and an active Secret Clearance with eligibility for TS/SCI.

San DiegoLast seen 4 days ago
$153,000 – $170,000 · Posted 7 days ago

Platform Operations Engineer to build and operate healthcare infrastructure on AWS EKS, managing container environments, CI/CD pipelines via GitHub Actions, and observability stacks with Datadog and CloudWatch. Responsibilities include supporting reliability through SLI/SLO tracking, on-call rotations, and incident response; maintaining GitOps workflows with ArgoCD; and writing infrastructure automation in Python.… Requires 3+ years SRE/DevOps experience, hands-on AWS (VPC, IAM, EKS, RDS), Terraform or CloudFormation, Kubernetes, and familiarity with observability tools.

San DiegoLast seen 5 days ago
Posted 8 days ago

Design, build, and maintain real-time data ingestion pipelines that reliably flow streaming data from multiple sources into a data platform while ensuring quality, observability, and scalability. Monitor data health, develop resilient pipelines with proper error-handling and backpressure strategies, and automate monitoring of real-time feeds to alert on timeliness, volume, and distribution issues.… Collaborate with data, platform, and software engineers to optimize pipeline configuration, perform root-cause analysis on data outages, and work with security teams on encryption and data classification. Requires 3+ years in data engineering, data operations, or DevOps roles; proficiency with Python and bash; and hands-on experience with tools like Kafka, NiFi, Spark Streaming, Snowflake, Elasticsearch, Grafana, and Prometheus.

San DiegoLast seen 6 days ago
$122,570 – $204,249 · Posted 8 days ago

The AVP Tech, Edge Security Engineer designs, builds, and supports highly available enterprise edge infrastructure and security controls using CDN platforms like Cloudflare, Akamai, and Fastly. The role requires architecting edge security standards, performing threat hunting and mitigation, and leading proof-of-concepts for emerging technologies including AI-driven solutions.… Responsibilities include collaborating with enterprise and security architecture teams, improving observability and security posture across edge environments, and participating in on-call incident response. Required experience spans 4+ years with edge infrastructure and CDN platforms, 4+ years with enterprise monitoring tools (Splunk, Datadog, Dynatrace, Grafana), global load balancing solutions, enterprise DNS, and infrastructure automation (Terraform, Ansible, Python, Go).

San DiegoLast seen 6 days ago
Posted 11 days ago

The Site Reliability Engineer owns the systems and practices that keep production services available, performant, and recoverable at scale, focusing on Kubernetes-based workloads, cloud infrastructure, observability, and incident response. The role requires designing and operating highly available infrastructure on AWS or GCP using Kubernetes, Terraform, and infrastructure-as-code; building observability across services using Prometheus, Grafana, OpenTelemetry, and centralized logging; and automating deployment, scaling, backup, and recovery workflows through CI/CD pipelines.… The engineer will define SLIs and SLOs, lead incident response and root-cause analysis, harden systems through access controls and disaster recovery testing, and partner with application teams to improve service design and operational readiness. The position requires 3–8 years of DevOps or platform engineering experience, hands-on Kubernetes and containerized services operation, strong Linux and networking fundamentals, proficiency with Terraform or equivalent infrastructure-as-code, observability implementation experience, and strong scripting or programming skills in Python, Go, or Bash.

San DiegoLast seen 9 days ago
Posted 11 days ago

The DevOps Engineer owns the design, deployment, and operation of production infrastructure at scale across AWS and Kubernetes, with responsibility for CI/CD pipelines, observability, and incident response. The role requires hands-on expertise in AWS services (EC2, EKS, VPC, IAM, S3, RDS), Kubernetes cluster operations, Terraform-based infrastructure-as-code, and monitoring tools like Prometheus and Grafana.… You will automate operational workflows with Python, Go, or Bash, lead incident investigations, and partner with software engineers and SREs to improve release velocity and system reliability. 3–8 years of DevOps, SRE, or platform engineering experience is expected.

San DiegoLast seen 9 days ago
Posted 12 days ago

Senior Network Engineer to design, build, and operate critical network infrastructure for cloud, AI/GPU, and high-performance computing environments. Responsibilities include managing multi-homed internet edge networks with BGP/eBGP, designing high-availability architectures with redundant firewalls and load balancing, deploying InfiniBand/GPU networks, automating configuration with Python/Ansible, and leading incident response.… Requires 7+ years of network engineering experience, strong BGP and TCP/IP expertise, hands-on work with Cisco/Arista/Fortinet platforms, and network automation skills.

San DiegoLast seen 10 days ago
$160,000 – $290,000 · Posted 14 days ago

Staff Engineer who will design, build, and operate the Forge Platform—a distributed-systems foundation serving autonomy, ML Ops, simulation, and application teams. You will own architecture and technical standards for workflow orchestration, asynchronous processing, event-driven systems, and long-running service workflows while remaining hands-on in implementation, production troubleshooting, and reliability improvement.… Required: strong production experience with distributed systems, cloud-native platforms, or backend infrastructure; fluency in Go and Python; deep understanding of failure handling, consistency, fault tolerance, and state management; and ability to turn recurring infrastructure needs into reusable platform capabilities. You will work across Kubernetes, service networking, observability, and data pipelines to enable downstream teams to move faster on mission-critical systems.

San DiegoLast seen 12 days ago
$160,000 – $290,000 · Posted 14 days ago

Staff Engineer responsible for designing, building, and operating the Forge Platform—a distributed-systems foundation serving autonomy, ML, simulation, and test teams. You will architect cloud-native backend infrastructure, establish reusable platform capabilities, define technical standards, and remain hands-on in production troubleshooting, performance analysis, and reliability improvement.… Required: deep expertise in production distributed systems, Go and Python, workflow orchestration, and the ability to translate complex architecture into usable interfaces across multiple downstream teams.

San DiegoLast seen 13 days ago
$120,001 – $160,000 · Posted 14 days ago

A Software Integration Engineer at SAIC will design, develop, and deploy a hybrid container/virtualization platform supporting enterprise-scale workloads across on-prem and cloud environments. The role encompasses infrastructure-as-code automation, Kubernetes cluster configuration, DevSecOps practices, and integration of core platform services (identity/SSO, PKI, DNS, logging/SIEM, monitoring, secrets management).… Responsibilities include automating verification/validation, establishing CI/CD pipelines and standard operating procedures, supporting GPU-enabled AI/ML workloads, and mentoring junior engineers on cloud-native best practices. The position requires 7–9+ years of experience with cloud platforms (AWS, Azure, GCP), containerization (Kubernetes, Docker), virtualization (VMware, Hyper-V, KVM), infrastructure-as-code tools (Terraform, Ansible, Puppet), and DevSecOps toolchains (ArgoCD, GitLab, Jenkins).

San DiegoLast seen 13 days ago
Posted 15 days ago

As a Staff Site Reliability Engineer at Altium (Renesas), you will ensure reliability, availability, and performance of large-scale SaaS cloud platforms through a combination of software engineering and systems administration. You will pioneer improvements in observability (logging, monitoring, APM), develop reliability frameworks, contribute to incident response and management, and drive automation and Infrastructure as Code initiatives across multiple regions.… The role requires 6+ years of SRE/DevOps experience in large-scale environments, 3+ years of software development (ideally .NET), strong knowledge of Kubernetes, AWS, microservices, and HA architecture, plus hands-on expertise with observability tools, CI/CD platforms, and IaC tools.

San DiegoLast seen 13 days ago
Posted 18 days ago

This is a Staff Site Reliability Engineer role focused on ensuring reliability, availability, and performance of Altium's large-scale SaaS cloud platforms. The role combines software engineering and systems administration, requiring 6+ years of SRE/DevOps experience and 3+ years of software development (preferably .NET).… Responsibilities include designing observability frameworks, automating operational tasks, managing incidents, implementing infrastructure-as-code, and collaborating with engineering teams on reliability best practices across AWS and Kubernetes environments.

San DiegoLast seen 16 days ago
Posted 22 days ago

Principal Database Architect responsible for administering, architecting, and optimizing enterprise Microsoft SQL Server database platforms across production, development, test, and disaster recovery environments at scale. The role involves designing highly available and secure database architectures, managing ETL/SSIS workflows, performance tuning, backup/recovery, replication, and disaster recovery processes.… Requires 10+ years of SQL Server administration in complex enterprise environments, strong T-SQL and data modeling expertise, and experience with reporting platforms (SSRS, SSAS, Power BI). Will partner with software engineering, QA, and infrastructure teams to support application requirements and evaluate emerging database technologies.

San DiegoLast seen 20 days ago
$149,800 – $262,200 · Posted 25 days ago

Staff Software Engineer on the AI Experience Framework (AIUX) team, responsible for full-stack architectural decisions in a JavaScript/component-driven environment powering ServiceNow's AI-first user interfaces. You will own code from design through delivery, architect modular reusable component systems, set framework-level technical direction for AI integration and state management, and serve as an escalation point for production issues across multiple teams.… The role requires 8+ years of OO language experience (Java, C++, C#, Go), advanced expertise in modern UI frameworks (React, Angular, Vue, Lit), relational databases, and data structures/algorithms/design patterns at scale.

San DiegoLast seen 23 days ago
$149,800 – $262,200 · Posted 26 days ago

Staff Software Engineer role owning full-stack architecture for ServiceNow's AI Experience Framework (AIUX), building modular Lit-based web components that power conversation-first applications. Responsible for end-to-end code delivery from design through production, setting technical direction for framework-level APIs, and serving as an escalation point for production issues across multiple teams.… Requires 8+ years of OO language experience (Java, C++, C#, Go), advanced JavaScript/modern UI framework expertise (Angular, React, Vue, Lit), and deep knowledge of data structures, algorithms, design patterns, and performance optimization. Expected to mentor colleagues, coordinate cross-org architecture decisions, and architect scalable component systems for extensibility.

San DiegoLast seen 23 days ago
$129,700 – $207,400 · Posted 26 days ago

Principal Database Architect responsible for administering, optimizing, and architecting enterprise SQL Server database platforms across production, development, test, and disaster recovery environments. Will monitor performance, manage backups/restores/migrations, develop and optimize SSIS packages for ETL and data integration, review database designs, and partner with engineering, QA, and infrastructure teams to ensure reliability, scalability, and security.… Requires 10+ years of Microsoft SQL Server administration in complex enterprise environments, strong T-SQL and relational database expertise, and hands-on experience with performance tuning, HA/DR, replication, and SSIS workflows.

San DiegoLast seen 24 days ago
$160,001 – $200,000 · Posted 27 days ago

Senior Platform Engineer responsible for designing, building, and maintaining secure, scalable cloud infrastructure on AWS GovCloud to support AI/ML workloads and autonomous systems for DoD missions. Core duties include container orchestration (Kubernetes/OpenShift), Infrastructure as Code (Terraform/Ansible/CloudFormation), DevSecOps pipeline collaboration, and ensuring compliance with DoD cybersecurity standards and STIG requirements.… Requires 14+ years of cloud engineering, DevOps, or platform engineering experience with active TS clearance at start and ability to obtain TS/SCI. Will mentor junior engineers and partner with AI/ML teams to deploy advanced analytics platforms.

San DiegoLast seen 2 days ago
$160,001 – $200,000 · Posted 27 days ago

Senior Platform Engineer responsible for designing, building, and maintaining secure, scalable cloud infrastructure in AWS GovCloud to support AI/ML workloads and autonomous systems for DoD. You will manage container orchestration (Kubernetes, OpenShift), implement Infrastructure as Code (Terraform, Ansible, CloudFormation), ensure DoD cybersecurity and STIG compliance, optimize cloud resources, and collaborate with DevSecOps teams on deployment pipelines.… The role requires 9–14+ years of experience depending on degree, an active TS clearance at hire, and ability to obtain TS/SCI, with hands-on troubleshooting, mentorship of junior engineers, and technical documentation responsibilities.

San DiegoLast seen 24 days ago
Posted 27 days ago

The Site Reliability Engineer owns the reliability, scalability, and operational readiness of production services running on AWS and Kubernetes. Responsibilities include designing highly available infrastructure with Terraform and managed AWS services, building CI/CD pipelines with GitHub Actions and Argo CD, defining SLOs and implementing observability with Prometheus and Grafana, leading incident response, and automating operational work with Python, Go, or Bash.… The role requires 3–8 years of hands-on SRE or DevOps experience, production Kubernetes expertise, strong AWS and infrastructure-as-code knowledge, and proficiency with observability and deployment strategies.

San DiegoLast seen 25 days ago
$120,001 – $160,000 · Posted 1 month ago

Software Integration Engineer responsible for designing, developing, and deploying a hybrid container/virtualization platform-as-code solution supporting enterprise-scale workloads across on-prem and cloud environments. The role involves integrating core platform services (identity, PKI, DNS, logging, monitoring, secrets management), configuring container runtimes and networking, and developing infrastructure-as-code automation.… Requires 9+ years of experience (or equivalent with advanced degree) with cloud platforms, Kubernetes, virtualization, DevSecOps practices, CI/CD tooling, and automation frameworks, plus the ability to mentor junior engineers and collaborate across security, platform, and application teams on compliance, incident response, and AI/ML workload enablement.

San DiegoLast seen 1 month ago
$233,000 – $350,000 · Posted 1 month ago

Senior Staff Engineer responsible for designing and building an MLOps platform that supports distributed AI training, reinforcement learning, and foundation model development at scale. You will architect Kubernetes-native infrastructure for GPU workloads, design self-service AI development workflows, manage the data and model lifecycle, and lead platform distribution across cloud, on-premises, and air-gapped environments.… The role requires deep expertise in modern AI frameworks (PyTorch, Hugging Face Transformers), distributed systems, GPU scheduling, and cloud-native infrastructure, with a focus on enabling researchers and engineers to move from experimentation to production rapidly.

San DiegoLast seen 1 month ago
$171,900 – $300,800 · Posted 1 month ago

Lead a team of full-stack software engineers building ServiceNow's AI Experience Framework (AIUX) — a platform delivering AI-first, conversation-first user interfaces through modular Lit-based web components. You'll manage product development, set technical direction for framework-level APIs, guide architecture decisions across the request path, and represent AIUX in cross-org discussions.… The role requires deep expertise in JavaScript and Java/C++/C#/Go, 10+ years of relevant technology experience, 5+ years managing core engineering teams, and the ability to solve complex problems spanning AI integration, performance at scale, and multi-team coordination.

San DiegoLast seen 1 month ago
$140,000 – $185,000 · Posted 1 month ago

Senior Full Stack Engineer owning end-to-end manufacturing software platform development across API, frontend, data model, and background jobs. You will design and evolve core domain models (work orders, engineering changes, effectivity, defect records) and ship production-quality features that technicians, supervisors, and quality teams rely on daily.… The role requires deep expertise in TypeScript, React, Node.js backends, and Postgres with strong proficiency in schema design, migrations, and transactional consistency. You'll partner with product, design, and manufacturing stakeholders to translate operational needs into maintainable software while raising standards on testing and observability.

San DiegoLast seen 17 days ago
Posted 1 month ago

DevOps Engineer to build, maintain, and optimize AWS cloud infrastructure and CI/CD deployment systems, working with ECS, EKS, and EC2 workloads. The role requires hands-on experience with Infrastructure as Code (Terraform), containerization (Docker/Kubernetes), monitoring tools (Grafana), and CI/CD pipelines, with growing ownership of services and systems.… You'll troubleshoot infrastructure and application issues, collaborate with engineering teams on deployment workflows, and apply AI tooling to improve infrastructure automation and operational efficiency. The position reports to a DevOps Manager and is based in San Diego on a hybrid schedule.

San DiegoLast seen 16 days ago