LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer
San DiegoLast seen today
Summary
Build and optimize scalable LLM inference platforms at Qualcomm's Cloud AI team, implementing advanced serving techniques like KV-cache management, speculative algorithms, and model optimization. Contribute to production serving frameworks (vLLM, SGLang, Triton, TGI) and work with customers on deployment solutions. Collaborate with compiler, firmware, and platform teams to drive efficient serving through autoscaling, load balancing, and routing. Requires deep understanding of transformer architectures, strong PyTorch and Python skills, computer architecture knowledge, and hands-on experience profiling and optimizing deep learning workloads.
Data ScienceDevOps / InfrastructureAIMachine LearningPythonAttention MechanismsAutoscalingComputer ArchitectureCudaData StructuresDeep LearningDistributed SystemsGenerative AIInference OptimizationKernel OptimizationKserveKv Cache ManagementLLM DLLM Large Language ModelsLmcacheLoad BalancingML AcceleratorsModel OptimizationMoeMooncakeOllamaParallel ProgrammingProfilingPyTorchSglangSlmSpeculative DecodingSystem DesignTgiTorch CompileTorchdynamoTransformersTriton Inference ServerTriton KernelsVllmVlm