← Back to results

speculative-decoding jobs in San Diego

$158,400 – $237,600 · Posted 1 month ago

Staff/Sr. Staff Software Engineer role focused on AI inference optimization on Snapdragon platforms, including model optimization, quantization, graph transformations, and runtime execution for LLMs, LVMs, and LMMs.… You will design and implement graph lowering and optimization techniques within ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK, working across ML algorithms, inference systems, and hardware integration. The role requires 6–8+ years of software development experience, 3+ years in AI/ML inference or model optimization, deep expertise in Python and C/C++, PyTorch/ONNX, and transformer architectures. You will mentor junior engineers, drive features end-to-end, and collaborate across ML Research, hardware, product, and QA teams.

San DiegoLast seen 18 days ago
$140,800 – $211,200 · Posted 1 month ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch.… You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen 20 days ago
$158,400 – $237,600 · Posted 1 month ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions.… Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen 1 day ago