← Back to results

agent-as-judge jobs in San Diego

Posted 28 days ago

Build and scale evaluation harnesses and automation to measure how well AI-powered software development tools perform. You'll develop reproducible benchmarking systems, validate evaluation approaches against human judgment, and analyze results across repeated runs to identify variance, failure patterns, and optimization opportunities.… The role requires strong software engineering expertise in automation and testing, proficiency in Python and at least one other language (Java, JavaScript), and solid knowledge of Git, Docker, CI/CD pipelines, and AI/LLM evaluation methodologies.

San DiegoLast seen 26 days ago