← Back to results

anomaly-detection jobs in San Diego

$92,400 – $148,800 · Posted 4 days ago

A Senior Site Reliability Engineer at SHEIN will own and operate mission-critical, large-scale distributed systems (Kubernetes, Kafka, Elasticsearch, Redis, APISIX, Nginx) running 24/7/365, participating in on-call rotations and driving fast incident response using AI-assisted log analysis and anomaly detection. The role requires strong software engineering expertise in Python or Go, deep Linux and networking knowledge, and hands-on experience with observability platforms (Prometheus, Grafana) and configuration management tools. Responsibilities include designing resilient monitoring and alerting infrastructure, automating operational workflows to eliminate toil, capacity planning, and collaborating with global teams to improve system reliability and performance.

San DiegoLast seen 2 days ago
$128,100 – $192,100 · Posted 5 days ago

Staff-level data engineer who designs, develops, and maintains scalable ETL/ELT pipelines using Databricks, PySpark, and Python for enterprise analytics and AI use cases. Builds curated data layers following Lakehouse Medallion architecture, develops Databricks-native applications (notebooks, dashboards, APIs), and optimizes Apache Spark workloads for performance and cost efficiency. Owns production data pipelines end-to-end, ensures data quality and reliability, and acts as a technical leader mentoring teams on data engineering best practices and modern Databricks capabilities.

San DiegoLast seen 4 days ago
Posted 6 days ago

Lead the product vision, strategy, and roadmap for diagnostics and debug analytics across pre-silicon and post-silicon semiconductor design phases. Own market analysis, PRD development, and technical requirements definition while serving as the subject-matter expert for DFT, silicon debug methodologies, and root-cause analysis. Partner with R&D, field teams, and data science groups to build diagnostic data pipelines, anomaly-detection models, and correlation engines that accelerate SoC development and yield ramp cycles. Engage directly with customers, executives, and cross-functional teams to translate complex debugging challenges into innovative analytics solutions and drive go-to-market strategy.

San DiegoLast seen 4 days ago
$177,300 – $265,900 · Posted 8 days ago

Design and implement Infrastructure as Code to automate provisioning, monitoring, and lifecycle management of NoSQL, Streaming, and Caching platforms (Cassandra, Aerospike, Kafka, Redis) across AWS and GCP. Build highly available, self-healing systems with automated failover and scaling, develop comprehensive observability solutions, and lead incident response for critical data platform issues. Drive automation-first practices and apply AI/ML approaches such as anomaly detection and predictive scaling to enhance reliability and reduce manual toil. Partner with engineering and platform teams to ensure resilient infrastructure supporting billions of transactions and millions of players globally.

San DiegoLast seen 6 days ago
Posted 10 days ago

As Senior Manager of AI Governance Technology at Teradata, you will own the end-to-end technical architecture of the company's AI Governance Operating System, including intake pipelines, compliance monitoring, and agentic workflows. You'll lead the engineering of sophisticated automation using tools like n8n and Claude APIs, architect the AI tool registry, and design real-time monitoring systems that anticipate governance challenges. You'll serve as the internal expert on Claude Enterprise, Microsoft Copilot, and Teradata AI Studio integrations, setting technical standards and specifications that guide the broader engineering team. This is a technical leadership role requiring deep expertise in enterprise AI governance, system architecture, and the ability to make complex engineering decisions with significant autonomy.

San DiegoLast seen 8 days ago
$108,000 – $180,000 · Posted 12 days ago

Staff Site Reliability Engineer at SHEIN responsible for operating and evolving large-scale, mission-critical production systems with 24/7/365 on-call participation. Design, build, and maintain observability solutions (metrics, logs, traces, alerting) with AI-powered anomaly detection; own and operate core open-source infrastructure (APISIX, Nginx, Kubernetes, Kafka, Elasticsearch, Redis, Consul, Etcd, Zookeeper). Automate operational workflows, reduce incident frequency and MTTR, and provide technical leadership across global engineering teams. Requires strong software engineering skills, deep Linux/networking/distributed systems expertise, and passion for solving problems at scale.

San DiegoLast seen 10 days ago