← Back to results

transformer-architectures jobs in San Diego

$140,800 – $211,200 · Posted 10 days ago

Design, develop, and optimize machine learning systems for production AI platforms, focusing on model development, inference optimization, and scalable ML infrastructure. Responsibilities include building ML pipelines and frameworks, optimizing model inference for latency and cost, integrating LLMs into APIs and microservices, and designing data pipelines for preprocessing and feature engineering.… The role requires strong software engineering fundamentals combined with deep ML expertise, working across PyTorch/TensorFlow, model-serving systems, distributed computing, and GPU environments.

San DiegoLast seen 8 days ago
$122,800 – $184,200 · Posted 10 days ago

Design, develop, and optimize production machine learning systems including model development, inference optimization, and scalable ML infrastructure. Build end-to-end ML pipelines from training through deployment, optimize model inference for latency and cost, integrate LLM/ML models into APIs and microservices, and engineer data pipelines for preprocessing and validation.… Requires strong software engineering fundamentals with deep ML expertise, proficiency in Python and systems languages, and experience with ML frameworks, model serving, and distributed computing.

San DiegoLast seen 8 days ago
$158,400 – $237,600 · Posted 1 month ago

Lead end-to-end model optimization and transformation for large language models, vision language models, diffusion, and multimodal models on Qualcomm inference accelerators. Architect PyTorch-based optimization strategies, drive graph capture and deployment using PyTorch/ONNX/torch.compile, and design fusion kernels using Triton or similar DSLs.… Partner with compiler, performance, and accuracy teams to co-design lowering strategies, optimize transformer-specific patterns (KVcache, decoding, long context), and scale distributed inference across multi-core and multi-device systems. Requires expert-level PyTorch proficiency, deep transformer architecture knowledge, hands-on experience with torch.compile/TorchDynamo, and strong foundation in ML accelerators and distributed systems.

San DiegoLast seen 1 month ago
$158,400 – $237,600 · Posted 1 month ago

Staff/Sr. Staff Software Engineer role focused on AI inference optimization on Snapdragon platforms, including model optimization, quantization, graph transformations, and runtime execution for LLMs, LVMs, and LMMs.… You will design and implement graph lowering and optimization techniques within ONNX Runtime, ExecuTorch, and Qualcomm AI Stack SDK, working across ML algorithms, inference systems, and hardware integration. The role requires 6–8+ years of software development experience, 3+ years in AI/ML inference or model optimization, deep expertise in Python and C/C++, PyTorch/ONNX, and transformer architectures. You will mentor junior engineers, drive features end-to-end, and collaborate across ML Research, hardware, product, and QA teams.

San DiegoLast seen 2 days ago
$178,400 – $267,600 · Posted 1 month ago

Qualcomm is seeking an AI Performance Engineer to optimize and deploy machine learning models for inference acceleration on cloud AI hardware. The role involves converting and optimizing language models, vision models, and diffusion models using PyTorch and ONNX, analyzing performance bottlenecks, and designing efficient kernels in Triton.… You will collaborate with compiler, firmware, and platform teams to map next-generation AI workloads onto hardware, working across the full product lifecycle from research to commercial deployment. The position requires deep expertise in transformer architectures, inference optimization techniques, computer architecture, ML accelerators, and distributed systems.

San DiegoLast seen 9 days ago
$140,800 – $211,200 · Posted 1 month ago

Design, develop, and optimize machine learning systems and models for production AI platforms, with focus on inference optimization, scalable ML infrastructure, and deployment. Responsibilities include building ML pipelines, optimizing model inference across hardware environments, integrating LLMs and models into APIs and microservices, designing data pipelines for ingestion and feature engineering, and collaborating cross-functionally on end-to-end ML solutions.… Requires strong software engineering fundamentals combined with deep ML expertise, proficiency in Python and at least one systems language (C++, Rust, or Go), and solid understanding of ML frameworks, transformer architectures, and model deployment systems.

San DiegoLast seen 1 month ago
Posted 1 month ago

Design and optimize video analysis, quality assessment, and encoding algorithms for hardware-accelerated video codec solutions (FPGA/ASIC). Develop C-models and firmware for real-time video processing, implementing algorithms for ROI detection, content understanding, frame-level rate control, and objective quality metrics.… Collaborate with infrastructure and platform teams to integrate algorithms into production VOD and live-streaming workflows, validate performance through A/B testing, and optimize end-to-end video quality. The role requires strong programming skills in C/C++ or Python, knowledge of video quality standards (VMAF, PSNR, SSIM), and ideally experience with deep learning frameworks, video codecs, or embedded real-time systems.

San DiegoLast seen 9 days ago