← Back to results

apache-spark jobs in San Diego

$122,600 – $177,900 · Posted 1 day ago

SHEIN is seeking a Senior Data Engineer to build and productionize GenAI/LLM solutions for data engineering workflows, including code assistance, metadata discovery, and knowledge access. The role focuses on developing RAG and agentic workflows, automating incident triage and log analysis, and integrating AI capabilities into internal developer tools and services. You'll own scoped projects from problem definition through production support, requiring 3+ years of hands-on production experience with strong Python/SQL fundamentals, proven GenAI/LLM application development, and practical knowledge of databases, APIs, data pipelines, and distributed systems.

San DiegoLast seen today
$140,000 – $190,000 · Posted 2 days ago

Founding Product Engineer to build a greenfield physical AI product at Netradyne, working across full-stack infrastructure, product UX, and agentic AI features. You'll design multi-tenant cloud and edge systems on AWS with real-time dashboards, compliance analytics, and AI-powered agent orchestration for safety operations. Expected to own end-to-end problems—from device to cloud to customer—with 7+ years of software engineering across backend (Go, Python, Java), cloud infrastructure (containers, CI/CD, distributed systems), and modern frontend (TypeScript/JavaScript). Direct customer collaboration and rapid iteration in the field required; fluency with AI-assisted development tools (Claude, Cursor) expected.

San DiegoLast seen today
$202,500 – $274,000 · Posted 3 days ago

Staff Data Engineer role at Credit Karma building petabyte-scale data infrastructure and streaming platforms on GCP. You will design and operate Kafka-based streaming infrastructure, build cloud-native data pipelines using Dataflow/Beam, Flink, and Spark, and create persistence frameworks across Spanner, MySQL, and BigQuery. The role requires 7+ years backend/data systems experience in JVM languages (Scala/Java), expertise with high-throughput distributed systems, and deep knowledge of streaming platforms, Apache Beam, CDC patterns, encryption, and data governance frameworks. You'll also integrate AI/ML and generative AI technologies including LLMs, RAG, semantic search, and knowledge graphs into the data platform.

San DiegoLast seen 1 day ago
$122,600 – $177,900 · Posted 3 days ago

Senior Data Engineer role focused on building and productionizing GenAI/LLM solutions for data engineering workflows. You will develop RAG and agentic workflows, automate incident triage and root-cause analysis, integrate AI capabilities into internal developer tools, and own projects from problem definition through production support. Required: 3+ years of production systems experience, strong Python/SQL, hands-on GenAI/LLM application building, and software engineering fundamentals including testing, CI/CD, and monitoring.

San DiegoLast seen 1 day ago
$128,100 – $192,100 · Posted 5 days ago

Staff-level data engineer who designs, develops, and maintains scalable ETL/ELT pipelines using Databricks, PySpark, and Python for enterprise analytics and AI use cases. Builds curated data layers following Lakehouse Medallion architecture, develops Databricks-native applications (notebooks, dashboards, APIs), and optimizes Apache Spark workloads for performance and cost efficiency. Owns production data pipelines end-to-end, ensures data quality and reliability, and acts as a technical leader mentoring teams on data engineering best practices and modern Databricks capabilities.

San DiegoLast seen 4 days ago
$134,500 – $265,100 · Posted 5 days ago

Forward Deployed Engineer who partners with enterprise clients to identify business needs and prototype production-grade GenAI/LLM solutions on AWS. Designs and builds AI-enabled agentic platforms, workflows, and scalable engineering patterns using AWS AI&Data services (Bedrock, Neptune, OpenSearch). Delivers production-quality code with strong testing, CI/CD, logging, and documentation practices while mentoring team members and translating complex business problems into technical AI solutions.

San DiegoLast seen 4 days ago
$198,200 – $297,400 · Posted 9 days ago

Staff Software Engineer leading the design, architecture, and operations of PlayStation's Real-Time Analytics Platform (RTAP), a large-scale distributed data system handling high-throughput, low-latency analytics and stream processing. The role involves architecting and evolving streaming and analytics platforms using Apache Flink, Spark, ClickHouse, and Druid; integrating batch and streaming workloads via lakehouse architectures; and embedding AI capabilities to automate engineering workflows and improve platform reliability. You will own critical platform capabilities end-to-end from design through production operations, mentor engineers, establish technical standards, and influence cross-functional technical strategy across PlayStation's data infrastructure.

San DiegoLast seen 7 days ago
$198,200 – $297,400 · Posted 9 days ago

Staff Software Engineer leading design and evolution of PlayStation's Real-Time Analytics Platform (RTAP), a large-scale distributed data system handling high-throughput, low-latency analytics and stream processing. Responsibilities include architecting stream processing pipelines, optimizing data services using Apache Flink, Spark, ClickHouse, and Druid, integrating batch and streaming workloads with lakehouse architectures, and building AI-powered capabilities for platform automation. The role requires deep expertise in distributed systems, real-time analytics, and mentoring engineering teams while establishing technical standards and best practices across the organization.

San DiegoLast seen 8 days ago
$198,200 – $297,400 · Posted 9 days ago

Staff-level engineer designing and operating large-scale distributed data platforms powering real-time analytics across PlayStation. You will lead the architecture and evolution of the Real-Time Analytics Platform (RTAP), own critical capabilities end-to-end from design through production operations, and mentor engineers. Key responsibilities include building and optimizing distributed services using Apache Flink, Spark Structured Streaming, ClickHouse, and Apache Druid; integrating batch and streaming workloads with lakehouse architectures; and establishing technical standards for distributed systems, stream processing, and data modeling.

San DiegoLast seen 7 days ago
Posted 9 days ago

S. Navy. a Developer, you will design, develop, and maintain scalable ETL/ELT pipelines and cloud-based data platforms that integrate enterprise systems and support AI/ML initiatives for the U.S. Navy. You will perform data engineering tasks including cleansing, transformation, and optimization of SQL and Apache Spark workloads on AWS, while also developing predictive models and generative AI solutions using Python-based technologies. You will build interactive dashboards and visualizations in Tableau or Qlik, translating technical findings into actionable insights for stakeholders. An active Secret clearance, U.S. citizenship, and a bachelor's degree in a technical field (or 4+ years of relevant professional experience) are required.

San DiegoLast seen 7 days ago
$110,000 – $130,000 · Posted 12 days ago

Design, build, and maintain enterprise data pipelines using Microsoft Fabric and Azure Synapse, handling ingestion, transformation, and orchestration from ERP and supply chain systems. Develop and manage Fabric Lakehouse and Warehouse architectures, apply data modeling best practices, and ensure data quality and reliability. Collaborate with analytics, BI, and IT teams to create governed, scalable analytical data assets that support Power BI and downstream consumers. Requires a bachelor's degree in a quantitative field, strong SQL and Python skills, and hands-on experience with Microsoft Fabric and Azure Synapse.

San DiegoLast seen 10 days ago