← Back to results

transformer-architectures jobs in San Diego

$158,400 – $237,600 · Posted 1 day ago

Staff/Sr. Staff Software Engineer role focused on AI inference optimization on Snapdragon platforms, including model optimization, quantization, graph transformations, and runtime execution for LLMs, LVMs, and LMMs. You will design and implement graph lowering and optimization techniques within ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK, working across ML algorithms, inference systems, and hardware integration. The role requires 6–8+ years of software development experience, 3+ years in AI/ML inference or model optimization, deep expertise in Python and C/C++, PyTorch/ONNX, and transformer architectures. You will mentor junior engineers, drive features end-to-end, and collaborate across ML Research, hardware, product, and QA teams.

San DiegoLast seen today
$178,400 – $267,600 · Posted 10 days ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton. You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 8 days ago
$140,800 – $211,200 · Posted 11 days ago

Design, develop, and optimize machine learning systems and models for production AI platforms, with focus on inference optimization, scalable ML infrastructure, and deployment. Responsibilities include building ML pipelines, optimizing model inference across hardware environments, integrating LLMs and models into APIs and microservices, designing data pipelines for ingestion and feature engineering, and collaborating cross-functionally on end-to-end ML solutions. Requires strong software engineering fundamentals combined with deep ML expertise, proficiency in Python and at least one systems language (C++, Rust, or Go), and solid understanding of ML frameworks, transformer architectures, and model deployment systems.

San DiegoLast seen 8 days ago
Posted 14 days ago

Design and optimize video analysis, quality assessment, and encoding algorithms for hardware-accelerated video codec solutions (FPGA/ASIC). Develop C-models and firmware for real-time video processing, implementing algorithms for ROI detection, content understanding, frame-level rate control, and objective quality metrics. Collaborate with infrastructure and platform teams to integrate algorithms into production VOD and live-streaming workflows, validate performance through A/B testing, and optimize end-to-end video quality. The role requires strong programming skills in C/C++ or Python, knowledge of video quality standards (VMAF, PSNR, SSIM), and ideally experience with deep learning frameworks, video codecs, or embedded real-time systems.

San DiegoLast seen 12 days ago