← Back to results

quantization jobs in San Diego

$169,500 – $254,300 · Posted 2 days ago

Qualcomm's GPU Research Team seeks GPU architects to design next-generation GPU hardware architectures for AI, ML, and GPGPU computing across mobile, Windows on Snapdragon, and data center platforms. The role involves collaborating with software and hardware teams to develop architectural solutions, optimize for workloads like large language models and vision models, and contribute to open-source GPU/ML projects.… Required skills include strong GPU architecture knowledge, proficiency with APIs (OpenCL, CUDA, Vulkan, Direct3D 12), and C/C++ programming; hands-on CUDA kernel optimization and familiarity with frameworks like llama.cpp or vLLM are highly valued.

San DiegoLast seen 1 day ago
$141,300 – $211,900 · Posted 3 days ago

Staff-level software engineer responsible for designing and executing subsystem integration testing (SSIT) strategy for Qualcomm's ML inference delegates portfolio. Build and maintain Python/PyTest-based test suites integrated into CI/CD pipelines, triage failures across the ML framework, QNN runtime, and HTP hardware stack, and serve as the technical liaison between development and QA/SIT teams.… Leverage AI-assisted tooling (Claude, Copilot) to accelerate test generation and failure analysis. Mentor junior engineers and establish team standards for test coverage and handoff documentation.

San DiegoLast seen 1 day ago
Posted 12 days ago

AI Research Scientist role focused on developing and optimizing large language models for on-device deployment on smartphones, IoT, and automotive systems. Responsibilities include researching and prototyping model compression, quantization, pruning, distillation, and inference optimization techniques.… The role involves collaboration with silicon, software, and product teams to design experiments, benchmark models, and translate research into production-ready solutions for 5G and edge AI applications.

San DiegoLast seen 10 days ago
$150,000 – $230,000 · Posted 13 days ago

Staff Engineer responsible for designing and building scalable test automation frameworks and CI/CD infrastructure for machine-learning systems at scale. The role requires expertise validating ML quality across the full lifecycle (data, training, evaluation, packaging, deployment, monitoring), GPU-accelerated workloads in Kubernetes, and integrated hardware-software systems.… You will develop automated evaluation suites, performance baselines, release thresholds, and observability solutions while mentoring teams on ML testing best practices and dependency/reproducibility standards.

San DiegoLast seen 11 days ago
$150,000 – $230,000 · Posted 13 days ago

Staff Engineer responsible for designing and implementing scalable test automation frameworks and infrastructure for machine learning systems across multiple engineering teams. The role requires expertise in validating ML pipelines across data, training, evaluation, deployment, and monitoring; testing GPU-accelerated workloads in Kubernetes; qualifying integrated hardware and software systems; and developing comprehensive integration and regression strategies.… Deep experience with Python automation, distributed systems testing, performance profiling, and observability is essential, along with strong system-design skills for complex, multi-tenant environments.

San DiegoLast seen 12 days ago
$186,700 – $280,100 · Posted 14 days ago

Qualcomm seeks a Computer Vision Engineer to develop machine learning algorithms for optical flow, depth estimation, visual tracking, SLAM, and 3D reconstruction on Snapdragon platforms. You will own end-to-end algorithm development—from training state-of-the-art neural networks and managing large datasets to optimizing and deploying models on-device for real-time performance.… The role requires deep expertise in ML model architecture, hardware/software optimization, and the ability to analyze performance bottlenecks across Qualcomm's hardware/software stack. You will serve as a technical leader, influencing system architecture, working with hardware teams on silicon design, and driving solutions from research through production.

San DiegoLast seen 12 days ago
$186,700 – $280,100 · Posted 15 days ago

Design and develop machine learning models for computer vision domains including optical flow, depth estimation, visual tracking, SLAM, and 3D reconstruction on Qualcomm Snapdragon platforms. Own end-to-end algorithm implementation from research through production deployment, including training pipeline development, dataset creation, model optimization, and on-device performance tuning.… Lead technical direction across projects, influence system-level architecture, and partner with hardware engineers on silicon design. Requires 6+ years of software/hardware/systems engineering experience (or equivalent with Master's/PhD) and demonstrated expertise in on-device ML deployment, neural network optimization, and computer vision algorithm development.

San DiegoLast seen 14 days ago
$200,000 – $250,000 · Posted 23 days ago

AppFolio is hiring a Staff Machine Learning Engineer to design, build, and operate their ML platform on AWS, supporting training, fine-tuning, inference, RAG, and cost optimization across the organization's AI initiatives. You'll partner with applied AI and research teams to productionize prototypes, maintain multi-provider LLM reliability (OpenAI, Google, Anthropic), and operate AI safety guardrails and authorization layers.… The role requires production-scale ML infrastructure experience on AWS (ECS, SageMaker, GPU fleets), deep knowledge of model serving and inference optimization, hands-on language model training, and demonstrated cost discipline across AI workloads.

San DiegoLast seen 22 days ago
$202,000 – $215,000 · Posted 24 days ago

Build production inference systems that turn ML research prototypes into deployed, performant components under strict latency, throughput, memory, and accuracy constraints. Own the full pipeline from model optimization through deployment: profile and tune tensor execution, write Rust/C++/CUDA where necessary, build evaluation machinery tied to product metrics, and maintain reproducible deployment contracts.… You'll collaborate with research teams to surface shipping risks early and translate algorithmic work into engineering reality with explicit performance budgets and regression gates.

San DiegoLast seen 22 days ago
Posted 26 days ago

Design, build, and operate production ML systems across the full lifecycle—from data preparation and model training through deployment, monitoring, and continuous improvement. Partner with data scientists, software engineers, and platform teams to turn research prototypes into reliable, scalable services.… Work across recommendation, ranking, forecasting, classification, NLP, and generative AI depending on product priorities. Balance model quality, inference latency, scalability, and operational resilience as equally important outcomes.

San DiegoLast seen 24 days ago
$122,500 – $213,200 · Posted 29 days ago

Design and architect next-generation mobile computer vision and deep learning systems for Qualcomm's AI accelerators and heterogeneous platforms. Define hardware-aware algorithm implementations, deep learning engine architectures, HW/SW partitioning strategies, and performance/power optimization across CPUs, GPUs, DSPs, NPUs, and dedicated accelerators.… Analyze neural network workloads, drive top-down architecture exploration, and collaborate with hardware teams to translate state-of-the-art CV and AI models into efficient mobile implementations.

San DiegoLast seen 27 days ago
$158,400 – $237,600 · Posted 1 month ago

Lead end-to-end model optimization and transformation for large language models, vision language models, diffusion, and multimodal models on Qualcomm inference accelerators. Architect PyTorch-based optimization strategies, drive graph capture and deployment using PyTorch/ONNX/torch.compile, and design fusion kernels using Triton or similar DSLs.… Partner with compiler, performance, and accuracy teams to co-design lowering strategies, optimize transformer-specific patterns (KVcache, decoding, long context), and scale distributed inference across multi-core and multi-device systems. Requires expert-level PyTorch proficiency, deep transformer architecture knowledge, hands-on experience with torch.compile/TorchDynamo, and strong foundation in ML accelerators and distributed systems.

San DiegoLast seen 1 month ago
$158,400 – $237,600 · Posted 1 month ago

Staff-level engineer responsible for designing and executing subsystem integration testing (SSIT) strategy for Qualcomm's Delegates ML inference framework. Develops and maintains Python/PyTest-based automated test suites integrated into CI pipelines, triages failures across the ML framework, QNN runtime, and HTP hardware stack, and owns the technical handoff criteria and traceability to downstream QA/SIT teams.… Leverages AI-assisted tooling (Claude Code, GitHub Copilot) to accelerate test generation and failure analysis. Mentors junior engineers and sets team standards for test coverage, debugging discipline, and documentation.

San DiegoLast seen 17 days ago
$158,400 – $237,600 · Posted 1 month ago

Staff-level software engineer responsible for subsystem integration testing (SSIT) of ML inference delegates on Qualcomm Snapdragon SoCs. Develops and maintains automated Python/PyTest test suites integrated into CI/CD pipelines, validates on-device behavior using hardware-in-the-loop infrastructure, and triages failures across the ML framework → QNN runtime → HTP hardware stack.… Acts as technical liaison between development and downstream QA/SIT teams, defining handoff criteria and owning test coverage strategy for the Delegates portfolio. Mentors junior engineers and sets team standards for test architecture, coverage discipline, and defect triage.

San DiegoLast seen 20 days ago
$158,400 – $237,600 · Posted 1 month ago

Staff/Sr. Staff Software Engineer role focused on AI inference optimization on Snapdragon platforms, including model optimization, quantization, graph transformations, and runtime execution for LLMs, LVMs, and LMMs.… You will design and implement graph lowering and optimization techniques within ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK, working across ML algorithms, inference systems, and hardware integration. The role requires 6–8+ years of software development experience, 3+ years in AI/ML inference or model optimization, deep expertise in Python and C/C++, PyTorch/ONNX, and transformer architectures. You will mentor junior engineers, drive features end-to-end, and collaborate across ML Research, hardware, product, and QA teams.

San DiegoLast seen 1 day ago
$140,800 – $211,200 · Posted 1 month ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch.… You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen 23 days ago
$140,800 – $211,200 · Posted 1 month ago

As a Senior Software Engineer focused on AI Tools, you will reauthor and optimize generative AI models (LLMs like Llama, Phi, Qwen, and multimodal models) for efficient execution on Qualcomm's on-device hardware. You'll translate hardware constraints into model-level transformations that preserve accuracy while enabling edge deployment, integrate inference acceleration techniques, and collaborate with compiler and quantization teams to move research prototypes into production.… The role requires deep implementation-level knowledge of generative AI architectures, strong Python proficiency in large typed codebases, and hands-on experience optimizing inference for resource-constrained environments.

San DiegoLast seen 27 days ago
$140,800 – $211,200 · Posted 1 month ago

Design, develop, and optimize machine learning systems for production AI platforms, including model development, inference optimization, and scalable ML infrastructure. Build training-to-deployment pipelines, optimize model serving for latency and cost, and integrate LLMs and generative AI models into microservices and APIs.… Develop data pipelines for ingestion, preprocessing, and feature engineering while collaborating cross-functionally with product, platform, and hardware teams to deliver end-to-end ML solutions.

San DiegoLast seen 1 month ago