← Back to results

apache-spark jobs in San Diego

$202,500 – $274,000 · Posted 1 month ago

Staff Data Engineer role at Credit Karma building petabyte-scale data infrastructure and streaming platforms on GCP. You will design and operate Kafka-based streaming infrastructure, build cloud-native data pipelines using Dataflow/Beam, Flink, and Spark, and create persistence frameworks across Spanner, MySQL, and BigQuery.… The role requires 7+ years backend/data systems experience in JVM languages (Scala/Java), expertise with high-throughput distributed systems, and deep knowledge of streaming platforms, Apache Beam, CDC patterns, encryption, and data governance frameworks. You'll also integrate AI/ML and generative AI technologies including LLMs, RAG, semantic search, and knowledge graphs into the data platform.

San DiegoLast seen 1 month ago
$122,600 – $177,900 · Posted 1 month ago

Senior Data Engineer role focused on building and productionizing GenAI/LLM solutions for data engineering workflows. You will develop RAG and agentic workflows, automate incident triage and root-cause analysis, integrate AI capabilities into internal developer tools, and own projects from problem definition through production support.… Required: 3+ years of production systems experience, strong Python/SQL, hands-on GenAI/LLM application building, and software engineering fundamentals including testing, CI/CD, and monitoring.

San DiegoLast seen 24 days ago
$128,100 – $192,100 · Posted 1 month ago

Staff-level data engineer who designs, develops, and maintains scalable ETL/ELT pipelines using Databricks, PySpark, and Python for enterprise analytics and AI use cases. Builds curated data layers following Lakehouse Medallion architecture, develops Databricks-native applications (notebooks, dashboards, APIs), and optimizes Apache Spark workloads for performance and cost efficiency.… Owns production data pipelines end-to-end, ensures data quality and reliability, and acts as a technical leader mentoring teams on data engineering best practices and modern Databricks capabilities.

San DiegoLast seen 5 days ago
$134,500 – $265,100 · Posted 1 month ago

Forward Deployed Engineer who partners with enterprise clients to identify business needs and prototype production-grade GenAI/LLM solutions on AWS. Designs and builds AI-enabled agentic platforms, workflows, and scalable engineering patterns using AWS AI&Data services (Bedrock, Neptune, OpenSearch).… Delivers production-quality code with strong testing, CI/CD, logging, and documentation practices while mentoring team members and translating complex business problems into technical AI solutions.

San DiegoLast seen 1 month ago
$198,200 – $297,400 · Posted 1 month ago

Staff Software Engineer leading the design, architecture, and operations of PlayStation's Real-Time Analytics Platform (RTAP), a large-scale distributed data system handling high-throughput, low-latency analytics and stream processing. The role involves architecting and evolving streaming and analytics platforms using Apache Flink, Spark, ClickHouse, and Druid; integrating batch and streaming workloads via lakehouse architectures; and embedding AI capabilities to automate engineering workflows and improve platform reliability.… You will own critical platform capabilities end-to-end from design through production operations, mentor engineers, establish technical standards, and influence cross-functional technical strategy across PlayStation's data infrastructure.

San DiegoLast seen 1 month ago
$198,200 – $297,400 · Posted 1 month ago

Staff Software Engineer leading design and evolution of PlayStation's Real-Time Analytics Platform (RTAP), a large-scale distributed data system handling high-throughput, low-latency analytics and stream processing. Responsibilities include architecting stream processing pipelines, optimizing data services using Apache Flink, Spark, ClickHouse, and Druid, integrating batch and streaming workloads with lakehouse architectures, and building AI-powered capabilities for platform automation.… The role requires deep expertise in distributed systems, real-time analytics, and mentoring engineering teams while establishing technical standards and best practices across the organization.

San DiegoLast seen 1 month ago
$198,200 – $297,400 · Posted 1 month ago

Staff-level engineer designing and operating large-scale distributed data platforms powering real-time analytics across PlayStation. You will lead the architecture and evolution of the Real-Time Analytics Platform (RTAP), own critical capabilities end-to-end from design through production operations, and mentor engineers.… Key responsibilities include building and optimizing distributed services using Apache Flink, Spark Structured Streaming, ClickHouse, and Apache Druid; integrating batch and streaming workloads with lakehouse architectures; and establishing technical standards for distributed systems, stream processing, and data modeling.

San DiegoLast seen 9 days ago
Posted 1 month ago

As a Data Developer, you will design, develop, and maintain scalable ETL/ELT pipelines and cloud-based data platforms that integrate enterprise systems and support AI/ML initiatives for the U.S. Navy.… You will perform data engineering tasks including cleansing, transformation, and optimization of SQL and Apache Spark workloads on AWS, while also developing predictive models and generative AI solutions using Python-based technologies. You will build interactive dashboards and visualizations in Tableau or Qlik, translating technical findings into actionable insights for stakeholders. An active Secret clearance, U.S. citizenship, and a bachelor's degree in a technical field (or 4+ years of relevant professional experience) are required.

San DiegoLast seen 1 month ago
$110,000 – $130,000 · Posted 1 month ago

Design, build, and maintain enterprise data pipelines using Microsoft Fabric and Azure Synapse, handling ingestion, transformation, and orchestration from ERP and supply chain systems. Develop and manage Fabric Lakehouse and Warehouse architectures, apply data modeling best practices, and ensure data quality and reliability.… Collaborate with analytics, BI, and IT teams to create governed, scalable analytical data assets that support Power BI and downstream consumers. Requires a bachelor's degree in a quantitative field, strong SQL and Python skills, and hands-on experience with Microsoft Fabric and Azure Synapse.

San DiegoLast seen 1 month ago