Don't ship what you can't measure: Evaluation-driven development for AI analytics agents
Building AI analytics agents is becoming democratised; tools like Snowflake Cortex Analyst make it easier than ever to deploy a text-to-SQL agent over your data models. However, trusting the output is still a challenge, and evaluating AI agents in analytics remains a bottleneck for most teams. As a rule of thumb: don't ship what you can't measure.This talk introduces evaluation-driven development for AI analytics agents: define what "correct" means before you build, measure every iteration, and only deploy when the evidence proves it's ready. Built from deploying Cortex Analyst agents leveraging modelled data in dbt across healthcare and e-commerce clients, this methodology is practical and repeatable.
You'll learn how to design golden question sets that test real business scenarios, build automated eval pipelines scoring across six dimensions, and track accuracy improvement across iterative semantic view enrichment.
Check out more sessions
- Breakout session
Governed by default: how data teams at Nordstrom turn dbt governance into an AI advantage
Priya Tanwar / NordstromNadine Bruxel / NordstromView session - Breakout session
AI-powered data development: How agentic SDLC & dbt-MCP transformed our data engineering workflow
Suhas Jangoan / ZendeskAyan Putatunda / ZendeskView session - Breakout session
No drift allowed: LangChain's context playbook with Hex
Emily Hawkins / LangChainLogan Cochran / LangChainView session
