Migaru AI builds an AI-native CRM that reads a sales team's email, calendar and calls to keep accounts, contacts and deals up to date automatically, drafting follow-ups and flagging decisions for a rep's approval rather than acting unsupervised. We're looking for an AI Evaluation Engineer to help us measure and improve the quality, accuracy and reliability of the AI systems that power the product, from record-keeping and summarization to agent-driven suggestions and drafts.
What you'll do
- Design and build evaluation frameworks and test suites for LLM-based features, including summarization, record extraction, agent reasoning, and generated drafts
- Define metrics for correctness, hallucination rate, relevance and usefulness, and track them across model and prompt changes
- Create and curate labeled datasets and golden sets to benchmark model and pipeline performance
- Analyze failure cases, identify root causes, and work with engineers to improve prompts, retrieval, and model choices
- Build tooling and dashboards that make evaluation results visible and actionable to the broader team
- Partner with product and engineering to decide what good output looks like for a given feature before and after it ships
What we're looking for
- Experience evaluating or testing machine learning or LLM-based systems in production
- Strong Python skills and comfort building data pipelines and tooling from scratch
- Familiarity with LLM evaluation techniques, such as human-in-the-loop review, model-graded evaluation, or benchmark construction
- Understanding of common failure modes in generative AI systems, including hallucination, drift, and prompt sensitivity
- Comfortable working with ambiguous, judgment-heavy problems and turning them into measurable criteria
- Clear written communication for documenting findings and recommending fixes
Nice to have
- Experience evaluating systems that read unstructured data such as email, call transcripts, or chat logs
- Background working on or alongside a CRM, sales, or customer-facing product
- Experience with retrieval-augmented generation or agentic workflows
