← Back to results

apache-spark jobs in San Diego

$62,000 – $141,000 · Posted 2 days ago

As a Data Engineer at Booz Allen Hamilton, you will design, develop, and deploy data pipelines and platforms that organize disparate data sources to support mission-critical analytics workloads. You'll work with ETL operations, relational and non-relational databases, and collaborate with analysts, developers, and data scientists in an agile environment.… Required skills include Python, SQL, Git, Linux/Windows automation, and experience with pipeline development or systems administration. The role emphasizes building scalable, maintainable data infrastructure that transforms raw data into actionable insights.

San DiegoLast seen 2 days ago
Posted 11 days ago

A forward-deployed engineer who builds and deploys AI/GenAI-powered solutions on Databricks and cloud platforms for forensic discovery and financial crime clients. The role requires hands-on expertise in Databricks, cloud services (AWS/Azure/GCP), and production AI/LLM systems, combined with the ability to lead client engagements, translate business requirements into technical architectures, and troubleshoot integration issues across enterprise environments.… You will design reusable assets, document workflows, and enable client teams to sustain delivered solutions while navigating security, compliance, and privacy reviews.

San DiegoLast seen 11 days ago
$130,000 – $150,000 · Posted 14 days ago

Lead hands-on modernization of legacy SSIS/SSRS ETL workloads into cloud-native Lakehouse pipelines using Databricks, dbt, Fivetran, and Airflow. Design and enforce standardized data frameworks (ingestion, transformation, curation, consumption), drive CI/CD and observability practices, and mentor engineers on modern data engineering patterns.… Evaluate emerging technologies (Delta Live Tables, Iceberg, streaming, AI-driven observability) through POCs and POVs, and leverage AI-assisted tools (Databricks Assistant, GitHub Copilot, Cursor AI) to accelerate development and reduce technical debt. Optimize Spark workloads and orchestration across Azure and GCP environments while partnering with architects, platform engineers, and business stakeholders.

San DiegoLast seen 13 days ago
$146,000 – $190,000 · Posted 22 days ago

Senior Data Engineer responsible for designing and implementing cloud data platforms (Databricks, Snowflake, BigQuery) and enterprise-scale data warehouses using Medallion Architecture. Will develop scalable ETL pipelines, optimize Spark jobs, integrate ML models into data workflows, and implement data governance and lineage tracking.… The role requires hands-on expertise in Python, PySpark, SQL, and AWS services, plus experience with AI/ML integration, DevOps practices (Terraform, CI/CD), and mentoring junior engineers and analysts.

San DiegoLast seen 20 days ago
$150,000 – $180,000 · Posted 22 days ago

Lead a Data Engineering team to establish modern DataOps and CI/CD practices, mentor engineers in PySpark, SQL, dbt, Fivetran, and cloud-native development, and design scalable data infrastructure across GCP, Azure, and AWS cloud data warehouses (BigQuery, Snowflake, Databricks). Manage technical delivery, collaborate across enterprise teams to develop end-to-end data solutions, and align data initiatives with business goals while remaining current on industry trends.

San DiegoLast seen 20 days ago
Posted 25 days ago

Design and review data engineering solution architectures on AWS and Databricks, aligning technical decisions with business goals and market trends. Lead enterprise-scale data migration and modernization programs, mentoring teams and driving initiatives from proposal through delivery.… Provide hands-on expertise in Databricks, PySpark, SQL, and AWS, with strong proficiency in CI/CD pipelines, data governance, and engineering best practices. Engage with clients and internal stakeholders to build productive relationships and ensure successful project execution.

San DiegoLast seen 23 days ago
Posted 25 days ago

This role is a Data Engineering Architect responsible for designing and reviewing enterprise-scale data architectures on Databricks and AWS, leading migration and modernization programs, and resolving complex technical challenges during build and deployment phases. The architect will provide strategic insights to development teams and clients, establish data governance frameworks, monitor system performance, and contribute to proposals and RFPs.… Required expertise includes hands-on proficiency in Databricks, PySpark, SQL, AWS, CI/CD pipelines, and data architecture best practices, with proven experience in enterprise-scale programs and strong technical leadership and stakeholder management skills.

San DiegoLast seen 23 days ago
$150,000 – $165,000 · Posted 26 days ago

Design, build, and maintain secure enterprise data platforms supporting AI/ML workloads, with responsibility for database optimization, CI/CD automation, and DoD compliance. Mentor junior engineers and provide technical leadership on architecture decisions.… Implement security controls aligned with NIST 800-53, RMF, and DISA STIG standards throughout the platform lifecycle. Develop ETL/ELT pipelines, support data lakes and lakehouses, and evaluate emerging data technologies for mission fit.

San DiegoLast seen 24 days ago
$140,000 – $190,000 · Posted 27 days ago

Founding Product Engineer at a safety/operations software company building AI-powered dashboards, compliance analytics, and agentic features across edge and cloud infrastructure. You'll own full-stack problems end-to-end—from AWS/container infrastructure and backend services (Go, Python, Java) through real-time frontends (TypeScript/JavaScript) to direct customer deployment and iteration.… Expected to ship fast, work across infrastructure, product, and AI agent layers, and demonstrate genuine range across the stack with 7+ years of software engineering experience.

San DiegoLast seen 25 days ago
$130,000 – $150,000 · Posted 27 days ago

Lead the modernization of legacy SSIS/SSRS data systems into cloud-native Lakehouse pipelines using Databricks, dbt, Fivetran, and Airflow. Design standardized frameworks for ingestion, transformation, and consumption layers while driving CI/CD, observability, and engineering best practices across Azure and GCP.… Evaluate emerging technologies (Delta Live Tables, Iceberg, streaming ingestion) through POCs and POVs, and leverage AI-assisted tools (Databricks Assistant, Cursor AI, GitHub Copilot) to accelerate development and reduce technical debt. Provide technical guidance to the team and translate business requirements into scalable data solutions.

San DiegoLast seen 25 days ago
$140,000 – $190,000 · Posted 28 days ago

Founding Product Engineer to build a greenfield Physical AI product for fleet safety and operations. The role spans full-stack development across cloud infrastructure (AWS, containers, CI/CD), multi-tenant backend systems with RBAC and SSO, real-time dashboards, compliance analytics, and AI-agentic features for alert triage and coaching workflows.… You'll own end-to-end problems from edge devices to cloud to user experience, shipping fast with direct customer feedback, and working across backend (Go, Python, Java), frontend (TypeScript/JavaScript), infrastructure, and AI integration. Requires 7+ years of software engineering with demonstrated range across multiple technical disciplines.

San DiegoLast seen 27 days ago
$135,375 – $135,375 · Posted 28 days ago

Principal-level individual contributor role leading AI/ML architecture and engineering strategy at scale. Requires 13+ years across ML, data engineering, and distributed systems, with 3+ years shipping production Generative/Agentic AI systems.… Hands-on expertise in RAG, vector databases, LLM optimization, agentic orchestration, and MLOps/LLMOps practice. Will set technical direction, own architecture decisions, drive platform initiatives, and raise engineering standards across the organization without direct people-management responsibility.

San DiegoLast seen 5 days ago
$187,000 – $187,000 · Posted 1 month ago

Lead a data engineering team responsible for building scalable data pipelines, trusted datasets, and reusable data products powering analytics, experimentation, and AI across Scribd. This role blends technical leadership with hands-on architecture guidance; you'll establish engineering standards, drive design reviews, mentor engineers, and partner cross-functionally to translate business needs into production-grade data solutions.… Requires 10+ years in data engineering or data platforms, 3+ years leading engineering teams, deep expertise in dimensional modeling and scalable data architectures, and strong SQL plus Python/Scala skills. You'll work with modern cloud data platforms like Databricks, Snowflake, and BigQuery, and distributed processing frameworks like Spark.

San DiegoLast seen 29 days ago
$85,000 – $100,000 · Posted 1 month ago

Design, build, and maintain Python and SQL-based data transformation workflows using Databricks and PySpark for large-scale processing. Analyze datasets to independently identify actionable insights, translate ambiguous data problems into clear analytic solutions, and ensure data quality and governance across pipelines.… Collaborate with stakeholders and cross-functional teams to integrate data outputs into downstream applications, troubleshoot pipeline issues, and document architecture and processes.

San DiegoLast seen 29 days ago
$85,000 – $100,000 · Posted 1 month ago

Design and build Python and SQL-based data transformation workflows using Databricks and PySpark for large-scale processing. Analyze unstructured datasets to independently identify actionable insights, translate ambiguous data problems into clear analytic solutions, and ensure data quality and governance across pipelines.… Troubleshoot pipeline issues, document processes and architecture, and partner with cross-functional teams to integrate data outputs into downstream applications and dashboards.

San DiegoLast seen 1 month ago
$132,000 – $132,000 · Posted 1 month ago

Senior Enterprise Data Engineer role requiring 5+ years designing and developing OLTP/OLAP databases, ETL/ELT pipelines using SSIS, Azure Data Factory, or Python, and enterprise reporting solutions with Power BI and SSRS. Must have advanced SQL Server expertise, strong data warehousing and dimensional modeling knowledge, and experience in Agile environments.… Preferred qualifications include Azure data platforms (Synapse, Data Lake, Fabric), Databricks, Apache Spark, Azure DevOps, and CI/CD practices.

San DiegoLast seen 1 month ago
Posted 1 month ago

Build and test agent behavior for an AI-driven service, including test case design and evaluation checks distinct from traditional software testing. Support data ingestion pipelines using AWS and BigData tools (PySpark, Airflow), collaborate with product and senior engineers to ship features, and contribute to production support (monitoring, alerting, incident response).… Requires 1–2 years of software engineering experience, a Master's in Computer Science or related field, proficiency in Python or another backend language, SQL/NoSQL fundamentals, basic REST API design, and hands-on experience with at least one LLM project (tool calling, chaining, retrieval, or multi-step coordination). Self-directed learner excited about AI tooling and focused on shipping end-to-end products.

San DiegoLast seen 1 month ago
$134,500 – $265,100 · Posted 1 month ago

Deloitte seeks a Forward Deployed Engineer to design, build, and deploy AI-enabled solutions and agentic platforms for government clients using Palantir technologies. The role requires hands-on experience building and operating GenAI/LLM-powered systems in production, contributing production-quality code with strong testing and CI/CD practices, and leading project workstreams to translate business problems into scalable AI solutions.… The engineer will prototype and deliver working solutions independently within a pod structure while mentoring teammates, balancing quality, safety, latency, cost, and model risk in architecture decisions.

San DiegoLast seen 1 month ago
$155,600 – $306,800 · Posted 1 month ago

Senior Forward Deployed Engineer at Deloitte GPS building AI-enabled solutions and agentic platforms on Databricks for enterprise government clients. Requires 7+ years software/data engineering experience, 5+ years deploying GenAI/LLM solutions in production, and 5+ years hands-on Databricks expertise across Lakehouse, Agent Bricks, Model Serving, and Genie.… Mentors junior engineers, designs scalable AI patterns with human-in-the-loop controls, delivers production-quality code with strong CI/CD and testing practices, and translates complex business problems into deployable AI solutions.

San DiegoLast seen 1 month ago
$124,000 – $190,000 · Posted 1 month ago

Design, build, and maintain scalable data pipelines and platforms that power high-volume, data-driven applications. Develop and optimize ETL/ELT processes for large, complex datasets using Databricks and Spark, and build backend data services in C# / .NET.… Lead technical design decisions, mentor engineers, and leverage AI-assisted coding tools to improve engineering efficiency. Requires 5+ years of data engineering experience, strong expertise with data pipelines and ETL frameworks, and hands-on experience with AWS-based data platforms.

San DiegoLast seen 14 days ago
$175,000 – $185,000 · Posted 1 month ago

Principal Engineer leading the design, build, and operation of an enterprise data platform serving 50+ source systems across R&D, Commercial, Manufacturing, and other business functions. The role combines hands-on data engineering (PySpark, T-SQL, Python, Data Factory, medallion architecture) with technical leadership of data engineers, vendor management, platform standards enforcement, and cross-functional stakeholder partnership.… Responsibilities span CI/CD pipeline design, data quality and reliability instrumentation, semantic layer development, data governance implementation, and delivery of new source-system integrations at scale.

San DiegoLast seen 1 month ago
$181,200 – $317,100 · Posted 1 month ago

As an IC5 Senior Staff Engineer, you will architect and deliver large-scale distributed data platform components centered on Kafka, Apache Iceberg, and Apache Spark. You will lead complex technical initiatives, design high-performance data ingestion pipelines, and define engineering best practices across the organization.… The role demands deep expertise in distributed systems, JVM performance tuning, stream processing, and full-stack Data Lake solutions, with hands-on delivery of production systems and mentorship of engineering teams.

San DiegoLast seen 1 month ago
$61,900 – $141,000 · Posted 1 month ago

As a Data Engineer at Booz Allen Hamilton, you will develop and deploy data pipelines and platforms that organize disparate data sources to yield actionable insights for mission-critical applications. You'll work with Python, SQL, or similar languages to build ETL operations, manage relational and non-relational databases, and support analytics workloads on cloud platforms.… The role requires 1+ years of data engineering, ETL, or pipeline development experience, proficiency with source control (GitHub/Atlassian), and Linux/Windows scripting. You'll collaborate with analysts, developers, and data scientists in an agile environment to design, develop, and maintain scalable data solutions.

San DiegoLast seen 1 month ago
$200,001 – $240,000 · Posted 1 month ago

Design, build, and maintain real-time data ingestion pipelines that reliably stream data from diverse sources into a production data platform. You will ensure data quality, observability, and scalability while monitoring pipeline health, building resilient systems with proper error handling and backpressure strategies, and automating monitoring and alerting for timeliness and data issues.… Partner with data engineers, platform engineers, and analytics teams to configure pipelines for reliability, perform root cause analysis on outages, and work with security teams on encryption and data governance. Proficiency required in Python and bash scripting, plus hands-on experience with tools like Kafka, NiFi, Flink, Spark, Snowflake, Grafana, and Prometheus.

San DiegoLast seen 1 month ago
$122,600 – $177,900 · Posted 1 month ago

SHEIN is seeking a Senior Data Engineer to build and productionize GenAI/LLM solutions for data engineering workflows, including code assistance, metadata discovery, and knowledge access. The role focuses on developing RAG and agentic workflows, automating incident triage and log analysis, and integrating AI capabilities into internal developer tools and services.… You'll own scoped projects from problem definition through production support, requiring 3+ years of hands-on production experience with strong Python/SQL fundamentals, proven GenAI/LLM application development, and practical knowledge of databases, APIs, data pipelines, and distributed systems.

San DiegoLast seen 1 month ago