Senior Staff Engineer, ML Ops (R4941)
ExpiredSan MateoLast seen 1 month ago
Summary
Senior Staff Engineer responsible for designing and building a Kubernetes-native MLOps platform that supports distributed AI training, reinforcement learning, and large-scale model development workflows. The role spans infrastructure architecture (GPU scheduling, compute optimization across cloud and on-premises), developer experience (self-service AI workflows), data and model lifecycle management, and cross-functional collaboration with ML researchers and autonomy teams. Required expertise includes Kubernetes, PyTorch, distributed training systems, Terraform, Helm, Python, and Golang; preferred experience includes Ray, KAI, Slurm, and observability tools like Prometheus and Grafana.
Data ScienceDevOps / InfrastructureAIGoKubernetesMachine LearningPythonTerraformAutonomyCloud NativeDistributed TrainingEdge DeploymentGpu SchedulingGrafanaHelmHugging Face TransformersInfrastructure As CodeKaiLinuxMlopsNetworkingObservabilityOpentelemetryPrometheusPyTorchRayReinforcement LearningRoboticsSecuritySlurmStorage