AI Evaluation Engineer
San DiegoLast seen 8 days ago
Summary
Build and scale evaluation tooling to measure how well AI-powered software development tools perform. Develop evaluation harnesses, automate benchmark runs, and ensure reproducibility through versioned processes with pinned dependencies and containerized environments. Validate evaluation approaches against human judgment, analyze results for variance and failure patterns, and work with engineering teams to improve automation workflows. This role requires strong software engineering foundations in automation and testing, hands-on experience with AI coding agents, and proficiency in Python, Java, or JavaScript.