AI Performance Engineer (Cloud AI Engineering), Sr | Staff | Sr. Staff
San DiegoLast seen today
Summary
Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton. You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.
Data ScienceDevOps / InfrastructureSoftware DevelopmentAIMachine LearningPythonAttention MechanismsComputer ArchitectureDiffusion ModelsDistributed SystemsInference OptimizationLinear AlgebraLLM OptimizationMachine Learning CompilersMemory OptimizationML AcceleratorsModel CompressionModel QuantizationNeural Network OperatorsOnnxPerformance ProfilingPyTorchTensor ParallelismTorch CompileTorchdynamoTransformer ArchitecturesTritonVlm OptimizationWorkload Mapping