← Back to results

attention-mechanisms jobs in San Diego

$158,400 – $237,600 · Posted 1 day ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions. Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen today
$139,000 – $183,000 · Posted 2 days ago

Senior Engineer, Machine Learning at Element Biosciences will design, develop, and optimize deep learning models (CNNs, Vision Transformers, U-Net) for biological image analysis and deploy them to production on AWS or imaging instruments. Responsibilities include building end-to-end ML pipelines from data ingestion through inference, applying advanced image processing and computer vision techniques to multimodal biological images, and analyzing single-cell and multiomic data. The role requires 5–7 years of experience with a Master's degree (or 0–3 years with a PhD) and hands-on proficiency in PyTorch, Python, and cloud deployment, with strong preference for experience in biomedical image modalities.

San DiegoLast seen today
$140,800 – $211,200 · Posted 3 days ago

As a Senior Software Engineer focused on AI Tools, you will reauthor and optimize generative AI models (LLMs like Llama, Phi, Qwen, and multimodal models) for efficient execution on Qualcomm's on-device hardware. You'll translate hardware constraints into model-level transformations that preserve accuracy while enabling edge deployment, integrate inference acceleration techniques, and collaborate with compiler and quantization teams to move research prototypes into production. The role requires deep implementation-level knowledge of generative AI architectures, strong Python proficiency in large typed codebases, and hands-on experience optimizing inference for resource-constrained environments.

San DiegoLast seen 1 day ago
$178,400 – $267,600 · Posted 6 days ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton. You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 4 days ago