← Back to results

model-compression jobs in San Diego

Posted 12 days ago

AI Research Scientist role focused on developing and optimizing large language models for on-device deployment on smartphones, IoT, and automotive systems. Responsibilities include researching and prototyping model compression, quantization, pruning, distillation, and inference optimization techniques.… The role involves collaboration with silicon, software, and product teams to design experiments, benchmark models, and translate research into production-ready solutions for 5G and edge AI applications.

San DiegoLast seen 10 days ago
$202,000 – $215,000 · Posted 24 days ago

Build production inference systems that turn ML research prototypes into deployed, performant components under strict latency, throughput, memory, and accuracy constraints. Own the full pipeline from model optimization through deployment: profile and tune tensor execution, write Rust/C++/CUDA where necessary, build evaluation machinery tied to product metrics, and maintain reproducible deployment contracts.… You'll collaborate with research teams to surface shipping risks early and translate algorithmic work into engineering reality with explicit performance budgets and regression gates.

San DiegoLast seen 22 days ago
$200,800 – $301,200 · Posted 1 month ago

Lead AI software strategy and architecture for Qualcomm's automotive platform, directing deployment of computer vision, perception systems, LLMs, and generative AI workloads. Drive architectural decisions across multiple products and customer programs while mentoring cross-functional teams.… Ensure solutions meet automotive safety, performance, and reliability standards. Requires 8+ years of software/systems engineering experience with deep expertise in AI/ML, C/C++, and automotive software environments.

San DiegoLast seen 29 days ago
$140,800 – $211,200 · Posted 1 month ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch.… You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen 23 days ago
$178,400 – $267,600 · Posted 1 month ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton.… You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 8 days ago