← Back to results

ai-evaluation jobs in San Diego

Posted 7 days ago

Senior Staff Inbound Product Manager responsible for owning the roadmap, vision, and requirements for conversational AI features within ServiceNow's AI platform. The role requires translating customer feedback into product improvements, guiding product strategy across multiple teams, and steering the development of AI features including disambiguation, clarification, document/multimodal experiences, slot filling, feedback intelligence, and fallback handling.… Success requires 12+ years of enterprise product management experience, deep knowledge of conversational AI, generative AI, and agentic systems, plus hands-on experience with AI evaluation, prompt engineering, retrieval systems, and knowledge bases. The person will create PRDs, define testing and quality requirements, support strategic customers through pilots and go-lives, and coordinate dependencies across engineering and design teams.

San DiegoLast seen 5 days ago
$60,000 – $148,500 · Posted 10 days ago

Build and scale evaluation tooling to measure how well AI-powered software development tools perform. Develop evaluation harnesses, automate benchmark runs, and ensure reproducibility through versioned processes with pinned dependencies and containerized environments.… Validate evaluation approaches against human judgment, analyze results for variance and failure patterns, and work with engineering teams to improve automation workflows. This role requires strong software engineering foundations in automation and testing, hands-on experience with AI coding agents, and proficiency in Python, Java, or JavaScript.

San DiegoLast seen 8 days ago
$128,900 – $219,100 · Posted 11 days ago

Design, build, test, and operate production ML and LLM-powered components for cybersecurity applications at scale. You will turn ambiguous requirements into working code, partner across product and security teams, and apply AI safety and guardrail practices.… The role requires solid software engineering fundamentals, hands-on experience building production ML/LLM applications (RAG, embeddings, agents), and the ability to take prototypes to reliable, maintainable systems. You'll work with distributed systems, cloud-native technologies, and graph/data plumbing while contributing to code reviews and raising team quality standards.

San DiegoLast seen 9 days ago
Posted 18 days ago

This Technical Lead role focuses on architecting, developing, and deploying production-grade Agentic AI solutions on AWS using .NET/C#. The ideal candidate will have hands-on expertise in LLMs, generative AI, RAG, prompt engineering, and agent frameworks (LangChain, LangGraph, Semantic Kernel, AutoGen), combined with strong .NET enterprise development and AWS cloud-native skills.… Responsibilities include taking AI solutions from POC to production, designing microservices and REST APIs, and ensuring AI evaluation, monitoring, observability, and governance. Technical leadership, stakeholder management, and system design fundamentals are essential.

San DiegoLast seen 16 days ago
Posted 28 days ago

A Forward Deployed AI Engineer who builds and deploys production AI/ML and Generative AI applications, particularly agentic workflows and RAG systems integrated with enterprise data and business systems. The role requires 6+ years of software/ML/AI engineering experience, advanced Python proficiency, hands-on expertise with LLMs, modern GenAI frameworks (LangChain, LangGraph, LlamaIndex), and cloud platforms (Azure, AWS, GCP).… You'll own solutions from prototype through production, working directly with product and business stakeholders to identify where AI creates value, troubleshoot integration issues, and drive continuous improvement.

San DiegoLast seen 7 days ago