← Back to results

inference-optimization jobs in San Diego

$158,400 – $237,600 · Posted 1 day ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions. Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen today
$140,800 – $211,200 · Posted 3 days ago

As a Senior Software Engineer focused on AI Tools, you will reauthor and optimize generative AI models (LLMs like Llama, Phi, Qwen, and multimodal models) for efficient execution on Qualcomm's on-device hardware. You'll translate hardware constraints into model-level transformations that preserve accuracy while enabling edge deployment, integrate inference acceleration techniques, and collaborate with compiler and quantization teams to move research prototypes into production. The role requires deep implementation-level knowledge of generative AI architectures, strong Python proficiency in large typed codebases, and hands-on experience optimizing inference for resource-constrained environments.

San DiegoLast seen 1 day ago
$178,400 – $267,600 · Posted 6 days ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton. You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 4 days ago
$141,200 – $278,300 · Posted 6 days ago

Lead the design and delivery of enterprise AI platforms and applications on Google Cloud, leveraging Vertex AI, Gemini, and cloud-native technologies. Design, fine-tune, and govern LLM solutions; build RAG and agentic systems; and define end-to-end architectures spanning data pipelines, feature engineering, model lifecycle, APIs, and MLOps/LLMOps. Architect cloud-native applications on GKE, Cloud Run, and managed services while implementing security, governance, and production-grade monitoring for AI/ML systems at scale.

San DiegoLast seen 4 days ago
$179,200 – $268,800 · Posted 7 days ago

As a Staff Product Data Manager for AI Platforms & Systems Solutions at Qualcomm, you will own the product strategy, definition, and lifecycle for AI platform products and specifications spanning model enablement, agentic systems, orchestration runtimes, and open protocols. You will develop comprehensive plans of record including schedules, budgets, and resources, directing products from conception through end of life while functioning as a central resource across systems, engineering, standards, quality, marketing, and partner teams. The role requires deep fluency in the Windows edge-AI stack, with technical depth across AI/ML models, agentic systems, orchestration engines, standards development, and hands-on experience delivering Windows-based AI products leveraging local accelerators (CPU, GPU, NPU). You will translate emerging model, agentic, and orchestration trends into differentiated product and specification roadmaps.

San DiegoLast seen 4 days ago
$111,300 – $166,900 · Posted 8 days ago

As a Senior Systems Engineer for Data Center AI at Qualcomm, you will research, develop, optimize, and validate AI/ML solutions that integrate Qualcomm's hardware accelerators with software and ecosystem to deliver high-performance inference in datacenter environments. You will design and deploy Gen AI and LLM applications, implement fine-tuning and distillation techniques, perform model benchmarking, and drive system-level architecture decisions. You will collaborate across functional teams to meet system-level requirements while staying current with advances in AI/ML models and hardware. Required skills include strong Python proficiency, expertise in ML frameworks (PyTorch, TensorFlow), deep understanding of ML deployment and system performance profiling, and experience with parallel computing.

San DiegoLast seen 6 days ago