← Back to results

vllm jobs in San Diego

$158,400 – $237,600 · Posted 1 day ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions. Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen today
$140,800 – $211,200 · Posted 7 days ago

Design, develop, and optimize machine learning systems and models for production AI platforms, with focus on inference optimization, scalable ML infrastructure, and deployment. Responsibilities include building ML pipelines, optimizing model inference across hardware environments, integrating LLMs and models into APIs and microservices, designing data pipelines for ingestion and feature engineering, and collaborating cross-functionally on end-to-end ML solutions. Requires strong software engineering fundamentals combined with deep ML expertise, proficiency in Python and at least one systems language (C++, Rust, or Go), and solid understanding of ML frameworks, transformer architectures, and model deployment systems.

San DiegoLast seen 4 days ago
$140,800 – $211,200 · Posted 8 days ago

Design, develop, and optimize machine learning systems for production AI platforms, including model development, inference optimization, and scalable ML infrastructure. Build training-to-deployment pipelines, optimize model serving for latency and cost, and integrate LLMs and generative AI models into microservices and APIs. Develop data pipelines for ingestion, preprocessing, and feature engineering while collaborating cross-functionally with product, platform, and hardware teams to deliver end-to-end ML solutions.

San DiegoLast seen 7 days ago
Posted 9 days ago

Design, build, and ship LLM-powered capabilities end to end—from prototyping and fine-tuning models to deploying production agents and retrieval systems. Own the full stack: prompt and context engineering, multi-step agent design with tool calling, RAG systems (embeddings, chunking, hybrid search, reranking), fine-tuning on multi-GPU with LoRA/QLoRA, evaluation and tracing infrastructure, and clean APIs for other engineers. Work in a secure, distributed environment where the platform runs on customer compute, cloud, or hybrid setups, requiring expertise with both commercial and self-hosted models.

San DiegoLast seen 7 days ago
$94,200 – $141,200 · Posted 10 days ago

As a Product Software Engineer, you will develop and integrate software for Qualcomm products spanning smartphones, computing devices, automotive infotainment, and IoT systems. You will own software integration tasks including component integration, build automation, version control management, and baseline releases. You will perform sanity testing on-target and in simulation environments, debug failures using tools like JTAG and ADB, automate test scenarios, and collaborate with cross-functional engineering teams on design and code reviews.

San DiegoLast seen 8 days ago