Perception Deployment Engineer - Model Deployment & Optimization
BostonLast seen 1 day ago
Summary
The Perception Deployment Engineer will optimize and deploy large-scale multi-modal foundation models, LLMs, and vision models to power- and thermal-constrained vehicle edge devices. Responsibilities include writing production C++ and CUDA code for real-time inference, implementing model compression (quantization, pruning, mixed-precision), architecting TensorRT compilation pipelines, and developing custom ML operators and CUDA kernels to achieve low-latency, deterministic execution. The role demands deep expertise in model optimization techniques, edge AI accelerators, and rigorous parity validation between PyTorch frameworks and compiled binaries.
Data ScienceC++Machine LearningPython3d Object Detection3d Occupancy NetworksAutonomous Driving PerceptionBevBf16CudaCustom ML OperatorsEdge DeploymentFlashattentionFoundation ModelsFp16Fp8Int8Kv Cache OptimizationLidarLinear AttentionLLM Large Language ModelsMixed Precision InferenceModel CompressionMulti Modal Sensor FusionOnnxPagedattentionPruningPtqPyTorchQatQuantizationRadarTensorrtTensorrt LLMTensorrt PluginsTorch CompileVisionVlaVlm