Data Scientist — Agent Evaluations & Quality

Clera

remote · Onsite · Full Time

Posted

Job description

About the Role This company is building an AI executive assistant that operates across email, calendars, meetings, and business software. As a Data Scientist — Agent Evaluations & Quality , you will own the measurement system that determines whether the assistant is genuinely improving in ambiguous, real-world environments. You'll partner directly with AI Agent Capabilities engineers to generate the evidence that shapes product decisions, model choices, and release quality. This is a high-ownership, deeply technical role at the intersection of applied data science, LLM evaluation, and product quality — ideal for someone who thrives on turning hard, open-ended quality questions into rigorous, actionable answers. What You'll Do Architect and maintain automated evaluation pipelines that measure agent quality across product surfaces. Translate agent capabilities into explicit pass, partial-pass, and failure criteria for complex multi-step tasks. Build representative gold datasets and regression suites covering real workflows, edge cases, and adversarial scenarios. Define meaningful metrics — task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability. Design deterministic and model-based graders, calibrate LLM-as-a-judge systems, and track grader agreement. Compare models, prompts, and implementations using rigorous offline experiments and production evidence. Analyze traces and production outcomes to identify root causes and build a practical failure taxonomy. Turn production failures into regression cases and continuously close gaps in evaluation coverage. Build dashboards and release-quality signals that make results actionable for engineering, product, and leadership. Recommend improvements to capability engineers and verify that fixes raise quality without unacceptable regressions. What We're Looking For Required 4+ years in Applied Data Science or Machine Learning roles, with a track record of building and delivering evaluation systems, automated data pipelines, or production ML infrastructure. Experience designing and implementing automated evaluation frameworks, success criteria, and regression suites for complex AI/ML or agentic systems. Production-grade proficiency in Python and SQL , with experience building and maintaining automated analytical pipelines on large datasets. Applied statistical and…

Apply for this job