← Back to results

pyspark jobs in San Diego

$175,000 – $185,000 · Posted 13 days ago

Principal Engineer leading the design, build, and operation of an enterprise data platform serving 50+ source systems across R&D, manufacturing, supply chain, and business functions. Responsible for hands-on architecture of medallion-layer pipelines using Microsoft Fabric, technical leadership and mentorship of data engineers, vendor management, data governance implementation, and platform reliability—balancing feature delivery with security, compliance (GxP), and operational excellence in a regulated biopharmaceutical environment.

San DiegoLast seen 11 days ago
Posted 16 days ago

Design and review data engineering solution architectures on Databricks and AWS, aligning technical decisions with business goals. Lead enterprise-scale data migration and modernization programs, providing architectural guidance and technical leadership across client and internal teams.… Demonstrate hands-on expertise in PySpark, SQL, data warehousing, CI/CD pipelines, and data governance frameworks. Mentor team members and foster a knowledge-sharing culture while driving smooth project execution and transition.

San DiegoLast seen 14 days ago
Posted 16 days ago

The role seeks a Databricks & AWS Data Engineering Architect to lead enterprise-scale data migration and modernization programs for a banking client. The architect will design and implement data solutions using Databricks, PySpark, SQL, and AWS, with responsibility for CI/CD pipelines, data governance frameworks, and data quality monitoring.… Required skills include hands-on expertise in data engineering, data warehousing, and data architecture, along with leadership capabilities to manage stakeholders and drive engineering best practices across the organization.

San DiegoLast seen 14 days ago
$146,000 – $190,000 · Posted 18 days ago

Senior Data Engineer responsible for designing and implementing cloud data platforms (Databricks, Snowflake, BigQuery) and enterprise-scale data warehouses using Medallion Architecture. Will develop scalable ETL pipelines, optimize Spark jobs, integrate ML models into data workflows, and implement data governance and lineage tracking.… The role requires hands-on expertise in Python, PySpark, SQL, and AWS services, plus experience with AI/ML integration, DevOps practices (Terraform, CI/CD), and mentoring junior engineers and analysts.

San DiegoLast seen 16 days ago
$150,000 – $180,000 · Posted 18 days ago

Lead a Data Engineering team to establish modern DataOps and CI/CD practices, mentor engineers in PySpark, SQL, dbt, Fivetran, and cloud-native development, and design scalable data infrastructure across GCP, Azure, and AWS cloud data warehouses (BigQuery, Snowflake, Databricks). Manage technical delivery, collaborate across enterprise teams to develop end-to-end data solutions, and align data initiatives with business goals while remaining current on industry trends.

San DiegoLast seen 16 days ago
Posted 18 days ago

Senior architect responsible for designing and leading enterprise-scale data platforms on Databricks and AWS, including cloud modernization, data migration, and lakehouse architecture. Requires deep expertise in PySpark, advanced SQL, Delta Lake, and data warehousing patterns, with accountability for technical strategy, solution design, and mentoring engineering teams.… Will drive CI/CD practices, data governance frameworks, data quality standards, and technical decision-making across complex data infrastructure initiatives. Expected to engage senior stakeholders, manage client relationships, and champion AI-powered engineering tools for productivity and automation.

San DiegoLast seen 16 days ago
$168,800 – $253,250 · Posted 19 days ago

Sr Manager leading a team of Platform Engineers/DataOps analysts to architect, deploy, and manage enterprise-scale Databricks Unity Catalog across dev/test/production environments. Responsibilities span platform governance, fine-grained access control (RBAC/ABAC), workspace automation via Terraform, CI/CD pipeline design (GitHub Actions, Azure DevOps, GitLab CI), FinOps optimization, and incident management for critical data infrastructure.… Requires 10+ years in data operations/platform engineering with 5+ years as a lead Databricks administrator, deep expertise in Unity Catalog governance, cloud security (IAM, VPC/VNet, compliance), and proficiency in Python, PySpark, and SQL.

San DiegoLast seen 17 days ago
Posted 21 days ago

Design and review data engineering solution architectures on AWS and Databricks, aligning technical decisions with business goals and market trends. Lead enterprise-scale data migration and modernization programs, mentoring teams and driving initiatives from proposal through delivery.… Provide hands-on expertise in Databricks, PySpark, SQL, and AWS, with strong proficiency in CI/CD pipelines, data governance, and engineering best practices. Engage with clients and internal stakeholders to build productive relationships and ensure successful project execution.

San DiegoLast seen 19 days ago
Posted 21 days ago

This role is a Data Engineering Architect responsible for designing and reviewing enterprise-scale data architectures on Databricks and AWS, leading migration and modernization programs, and resolving complex technical challenges during build and deployment phases. The architect will provide strategic insights to development teams and clients, establish data governance frameworks, monitor system performance, and contribute to proposals and RFPs.… Required expertise includes hands-on proficiency in Databricks, PySpark, SQL, AWS, CI/CD pipelines, and data architecture best practices, with proven experience in enterprise-scale programs and strong technical leadership and stakeholder management skills.

San DiegoLast seen 19 days ago
$140,000 – $190,000 · Posted 23 days ago

Founding Product Engineer at a safety/operations software company building AI-powered dashboards, compliance analytics, and agentic features across edge and cloud infrastructure. You'll own full-stack problems end-to-end—from AWS/container infrastructure and backend services (Go, Python, Java) through real-time frontends (TypeScript/JavaScript) to direct customer deployment and iteration.… Expected to ship fast, work across infrastructure, product, and AI agent layers, and demonstrate genuine range across the stack with 7+ years of software engineering experience.

San DiegoLast seen 21 days ago
$140,000 – $190,000 · Posted 24 days ago

Founding Product Engineer to build a greenfield Physical AI product for fleet safety and operations. The role spans full-stack development across cloud infrastructure (AWS, containers, CI/CD), multi-tenant backend systems with RBAC and SSO, real-time dashboards, compliance analytics, and AI-agentic features for alert triage and coaching workflows.… You'll own end-to-end problems from edge devices to cloud to user experience, shipping fast with direct customer feedback, and working across backend (Go, Python, Java), frontend (TypeScript/JavaScript), infrastructure, and AI integration. Requires 7+ years of software engineering with demonstrated range across multiple technical disciplines.

San DiegoLast seen 23 days ago
$135,375 – $135,375 · Posted 24 days ago

Principal-level individual contributor role leading AI/ML architecture and engineering strategy at scale. Requires 13+ years across ML, data engineering, and distributed systems, with 3+ years shipping production Generative/Agentic AI systems.… Hands-on expertise in RAG, vector databases, LLM optimization, agentic orchestration, and MLOps/LLMOps practice. Will set technical direction, own architecture decisions, drive platform initiatives, and raise engineering standards across the organization without direct people-management responsibility.

San DiegoLast seen 1 day ago
$150,000 – $180,000 · Posted 24 days ago

Lead a Data Engineering team to establish modern DataOps and CI/CD practices while providing technical mentorship in PySpark, SQL, dbt, and Fivetran. Design and optimize scalable, reliable data infrastructure across cloud and on-premise ecosystems (AWS, GCP, Azure) using Databricks, BigQuery, and Snowflake.… Manage end-to-end data sourcing, processing, integration, and delivery projects to meet business objectives and predetermined success criteria. Require 10+ years managing Data Engineering teams and 7+ years working with cloud-native data platforms and advanced SQL/Python/PySpark expertise.

San DiegoLast seen 22 days ago
$85,000 – $100,000 · Posted 28 days ago

Design, build, and maintain Python and SQL-based data transformation workflows using Databricks and PySpark for large-scale processing. Analyze datasets to independently identify actionable insights, translate ambiguous data problems into clear analytic solutions, and ensure data quality and governance across pipelines.… Collaborate with stakeholders and cross-functional teams to integrate data outputs into downstream applications, troubleshoot pipeline issues, and document architecture and processes.

San DiegoLast seen 25 days ago
$85,000 – $100,000 · Posted 28 days ago

Design and build Python and SQL-based data transformation workflows using Databricks and PySpark for large-scale processing. Analyze unstructured datasets to independently identify actionable insights, translate ambiguous data problems into clear analytic solutions, and ensure data quality and governance across pipelines.… Troubleshoot pipeline issues, document processes and architecture, and partner with cross-functional teams to integrate data outputs into downstream applications and dashboards.

San DiegoLast seen 26 days ago
Posted 1 month ago

Build and test agent behavior for an AI-driven service, including test case design and evaluation checks distinct from traditional software testing. Support data ingestion pipelines using AWS and BigData tools (PySpark, Airflow), collaborate with product and senior engineers to ship features, and contribute to production support (monitoring, alerting, incident response).… Requires 1–2 years of software engineering experience, a Master's in Computer Science or related field, proficiency in Python or another backend language, SQL/NoSQL fundamentals, basic REST API design, and hands-on experience with at least one LLM project (tool calling, chaining, retrieval, or multi-step coordination). Self-directed learner excited about AI tooling and focused on shipping end-to-end products.

San DiegoLast seen 28 days ago
$175,000 – $185,000 · Posted 1 month ago

Principal Engineer leading the design, build, and operation of an enterprise data platform serving 50+ source systems across R&D, Commercial, Manufacturing, and other business functions. The role combines hands-on data engineering (PySpark, T-SQL, Python, Data Factory, medallion architecture) with technical leadership of data engineers, vendor management, platform standards enforcement, and cross-functional stakeholder partnership.… Responsibilities span CI/CD pipeline design, data quality and reliability instrumentation, semantic layer development, data governance implementation, and delivery of new source-system integrations at scale.

San DiegoLast seen 1 month ago
$140,000 – $190,000 · Posted 1 month ago

Founding Product Engineer to build a greenfield physical AI product at Netradyne, working across full-stack infrastructure, product UX, and agentic AI features. You'll design multi-tenant cloud and edge systems on AWS with real-time dashboards, compliance analytics, and AI-powered agent orchestration for safety operations.… Expected to own end-to-end problems—from device to cloud to customer—with 7+ years of software engineering across backend (Go, Python, Java), cloud infrastructure (containers, CI/CD, distributed systems), and modern frontend (TypeScript/JavaScript). Direct customer collaboration and rapid iteration in the field required; fluency with AI-assisted development tools (Claude, Cursor) expected.

San DiegoLast seen 1 month ago
$128,100 – $192,100 · Posted 1 month ago

Staff-level data engineer who designs, develops, and maintains scalable ETL/ELT pipelines using Databricks, PySpark, and Python for enterprise analytics and AI use cases. Builds curated data layers following Lakehouse Medallion architecture, develops Databricks-native applications (notebooks, dashboards, APIs), and optimizes Apache Spark workloads for performance and cost efficiency.… Owns production data pipelines end-to-end, ensures data quality and reliability, and acts as a technical leader mentoring teams on data engineering best practices and modern Databricks capabilities.

San DiegoLast seen 1 day ago
Posted 1 month ago

Own and deliver small to medium machine learning system components from design through production deployment, including building data pipelines, training and evaluating models, and implementing MLOps monitoring. Write high-quality Python code to translate technical requirements into maintainable solutions, working with frameworks like scikit-learn, TensorFlow, PyTorch, and HuggingFace.… Design and deploy ML models as microservices, APIs, batch jobs, or streaming components on AWS, with responsibility for model performance metrics, data drift detection, and retraining triggers. Collaborate across Data Engineers, Software Engineers, Data Scientists, and product stakeholders to deliver project objectives.

San DiegoLast seen 1 month ago
Posted 1 month ago

As a Data Developer, you will design, develop, and maintain scalable ETL/ELT pipelines and cloud-based data platforms that integrate enterprise systems and support AI/ML initiatives for the U.S. Navy.… You will perform data engineering tasks including cleansing, transformation, and optimization of SQL and Apache Spark workloads on AWS, while also developing predictive models and generative AI solutions using Python-based technologies. You will build interactive dashboards and visualizations in Tableau or Qlik, translating technical findings into actionable insights for stakeholders. An active Secret clearance, U.S. citizenship, and a bachelor's degree in a technical field (or 4+ years of relevant professional experience) are required.

San DiegoLast seen 1 month ago