Build and scale evaluation harnesses and automation for AI-powered software development tools, turning real engineering artifacts into repeatable benchmark tasks. Develop versioned, reproducible processes with containerized environments to measure AI tool performance across quality, productivity, and efficiency metrics.… Validate evaluation approaches against human judgment, analyze results for variance and failure patterns, and work with engineering and data teams to improve tooling and document findings.