The Perception Deployment Engineer will optimize and deploy large-scale multi-modal foundation models, LLMs, and vision models to power- and thermal-constrained vehicle edge devices. Responsibilities include writing production C++ and CUDA code for real-time inference, implementing model compression (quantization, pruning, mixed-precision), architecting TensorRT compilation pipelines, and developing custom ML operators and CUDA kernels to achieve low-latency, deterministic execution.… The role demands deep expertise in model optimization techniques, edge AI accelerators, and rigorous parity validation between PyTorch frameworks and compiled binaries.