← Back to results

model-compression jobs in San Diego

$140,800 – $211,200 · Posted 2 days ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch. You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen today
$178,400 – $267,600 · Posted 9 days ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton. You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 7 days ago