← Back to results

kv-cache-management jobs in San Diego

$122,800 – $184,200 · Posted 6 days ago

Design, develop, and optimize production machine learning systems including model development, inference optimization, and scalable ML infrastructure. Build end-to-end ML pipelines from training through deployment, optimize model inference for latency and cost, integrate LLM/ML models into APIs and microservices, and engineer data pipelines for preprocessing and validation.… Requires strong software engineering fundamentals with deep ML expertise, proficiency in Python and systems languages, and experience with ML frameworks, model serving, and distributed computing.

San DiegoLast seen 4 days ago
$158,400 – $237,600 · Posted 1 month ago

Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions.… Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.

San DiegoLast seen 1 day ago