← Back to results

kernel-fusion jobs in San Diego

$202,000 – $215,000 · Posted 17 days ago

Build production inference systems that turn ML research prototypes into deployed, performant components under strict latency, throughput, memory, and accuracy constraints. Own the full pipeline from model optimization through deployment: profile and tune tensor execution, write Rust/C++/CUDA where necessary, build evaluation machinery tied to product metrics, and maintain reproducible deployment contracts.… You'll collaborate with research teams to surface shipping risks early and translate algorithmic work into engineering reality with explicit performance budgets and regression gates.

San DiegoLast seen 15 days ago
$158,400 – $237,600 · Posted 26 days ago

Lead end-to-end model optimization and transformation for large language models, vision language models, diffusion, and multimodal models on Qualcomm inference accelerators. Architect PyTorch-based optimization strategies, drive graph capture and deployment using PyTorch/ONNX/torch.compile, and design fusion kernels using Triton or similar DSLs.… Partner with compiler, performance, and accuracy teams to co-design lowering strategies, optimize transformer-specific patterns (KVcache, decoding, long context), and scale distributed inference across multi-core and multi-device systems. Requires expert-level PyTorch proficiency, deep transformer architecture knowledge, hands-on experience with torch.compile/TorchDynamo, and strong foundation in ML accelerators and distributed systems.

San DiegoLast seen 24 days ago