Design, build, and ship LLM-powered capabilities end to end—from prototyping and fine-tuning models to deploying production agents and retrieval systems. Own the full stack: prompt and context engineering, multi-step agent design with tool calling, RAG systems (embeddings, chunking, hybrid search, reranking), fine-tuning on multi-GPU with LoRA/QLoRA, evaluation and tracing infrastructure, and clean APIs for other engineers. Work in a secure, distributed environment where the platform runs on customer compute, cloud, or hybrid setups, requiring expertise with both commercial and self-hosted models.