← Back to results

attention-mechanisms jobs in San Diego

$129,500 – $194,300 · Posted 28 days ago

Qualcomm seeks an ML/computer vision engineer to develop and optimize AI-accelerated vision algorithms for edge devices (mobile, automotive, robotics, AR/VR). You will design algorithms for depth estimation, optical flow, video super-resolution, 3D reconstruction, and SLAM, then optimize deep learning models for resource-constrained SoCs.… The role spans algorithm research, model architecture design, performance profiling across memory/compute/power/bandwidth, and cross-functional collaboration with hardware and systems teams to ship production models for key customer segments.

San DiegoLast seen 25 days ago
$158,400 – $237,600 · Posted 1 month ago

Lead end-to-end model optimization and transformation for large language models, vision language models, diffusion, and multimodal models on Qualcomm inference accelerators. Architect PyTorch-based optimization strategies, drive graph capture and deployment using PyTorch/ONNX/torch.compile, and design fusion kernels using Triton or similar DSLs.… Partner with compiler, performance, and accuracy teams to co-design lowering strategies, optimize transformer-specific patterns (KVcache, decoding, long context), and scale distributed inference across multi-core and multi-device systems. Requires expert-level PyTorch proficiency, deep transformer architecture knowledge, hands-on experience with torch.compile/TorchDynamo, and strong foundation in ML accelerators and distributed systems.

San DiegoLast seen 29 days ago
$158,400 – $237,600 · Posted 1 month ago

Staff/Sr. Staff Software Engineer role focused on AI inference optimization on Snapdragon platforms, including model optimization, quantization, graph transformations, and runtime execution for LLMs, LVMs, and LMMs.… You will design and implement graph lowering and optimization techniques within ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK, working across ML algorithms, inference systems, and hardware integration. The role requires 6–8+ years of software development experience, 3+ years in AI/ML inference or model optimization, deep expertise in Python and C/C++, PyTorch/ONNX, and transformer architectures. You will mentor junior engineers, drive features end-to-end, and collaborate across ML Research, hardware, product, and QA teams.

San DiegoLast seen 19 days ago
$140,800 – $211,200 · Posted 1 month ago

Qualcomm AI Research seeks a Senior AI Research Quantization Engineer to develop algorithms for efficient generative AI, LLMs, and multimodal models optimized for on-device deployment. The role focuses on advanced quantization techniques, model compression, inference optimization (batching, KV caching, speculative decoding), and system prototyping using Python and PyTorch.… You will collaborate across hardware, software, and systems teams to enable state-of-the-art models to run on power- and memory-constrained devices including smartphones, autonomous vehicles, robotics, and IoT platforms. A Bachelor's degree plus 2+ years of related engineering experience (or Master's with 1+ year, or PhD) is required.

San DiegoLast seen 21 days ago
$158,400 – $237,600 · Posted 1 month ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions.… Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen 2 days ago
$139,000 – $183,000 · Posted 1 month ago

Senior Engineer, Machine Learning at Element Biosciences will design, develop, and optimize deep learning models (CNNs, Vision Transformers, U-Net) for biological image analysis and deploy them to production on AWS or imaging instruments. Responsibilities include building end-to-end ML pipelines from data ingestion through inference, applying advanced image processing and computer vision techniques to multimodal biological images, and analyzing single-cell and multiomic data.… The role requires 5–7 years of experience with a Master's degree (or 0–3 years with a PhD) and hands-on proficiency in PyTorch, Python, and cloud deployment, with strong preference for experience in biomedical image modalities.

San DiegoLast seen 24 days ago
$140,800 – $211,200 · Posted 1 month ago

As a Senior Software Engineer focused on AI Tools, you will reauthor and optimize generative AI models (LLMs like Llama, Phi, Qwen, and multimodal models) for efficient execution on Qualcomm's on-device hardware. You'll translate hardware constraints into model-level transformations that preserve accuracy while enabling edge deployment, integrate inference acceleration techniques, and collaborate with compiler and quantization teams to move research prototypes into production.… The role requires deep implementation-level knowledge of generative AI architectures, strong Python proficiency in large typed codebases, and hands-on experience optimizing inference for resource-constrained environments.

San DiegoLast seen 25 days ago
$178,400 – $267,600 · Posted 1 month ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton.… You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 6 days ago