← Back to results

quantization jobs in San Diego

$140,800 – $211,200 · Posted 2 days ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch. You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen today
$140,800 – $211,200 · Posted 6 days ago

As a Senior Software Engineer focused on AI Tools, you will reauthor and optimize generative AI models (LLMs like Llama, Phi, Qwen, and multimodal models) for efficient execution on Qualcomm's on-device hardware. You'll translate hardware constraints into model-level transformations that preserve accuracy while enabling edge deployment, integrate inference acceleration techniques, and collaborate with compiler and quantization teams to move research prototypes into production. The role requires deep implementation-level knowledge of generative AI architectures, strong Python proficiency in large typed codebases, and hands-on experience optimizing inference for resource-constrained environments.

San DiegoLast seen 4 days ago
$140,800 – $211,200 · Posted 11 days ago

Design, develop, and optimize machine learning systems for production AI platforms, including model development, inference optimization, and scalable ML infrastructure. Build training-to-deployment pipelines, optimize model serving for latency and cost, and integrate LLMs and generative AI models into microservices and APIs. Develop data pipelines for ingestion, preprocessing, and feature engineering while collaborating cross-functionally with product, platform, and hardware teams to deliver end-to-end ML solutions.

San DiegoLast seen 10 days ago