Don't ship what you can't measure: Evaluation-driven development for AI analytics agents
Building AI analytics agents is becoming democratised; tools like Snowflake Cortex Analyst make it easier than ever to deploy a text-to-SQL agent over your data models. However, trusting the output is still a challenge, and evaluating AI agents in analytics remains a bottleneck for most teams. As a rule of thumb: don't ship what you can't measure.This talk introduces evaluation-driven development for AI analytics agents: define what "correct" means before you build, measure every iteration, and only deploy when the evidence proves it's ready. Built from deploying Cortex Analyst agents leveraging modelled data in dbt across healthcare and e-commerce clients, this methodology is practical and repeatable.
You'll learn how to design golden question sets that test real business scenarios, build automated eval pipelines scoring across six dimensions, and track accuracy improvement across iterative semantic view enrichment.
Check out more sessions
- Breakout session
Revolutionizing the staging layer: how we automated model generation for 700+ sources
Cathy Huang / WebstaurantstoreView session - Breakout session
A nonprofit's journey: we walked so dbt could run
Kensly Alexander / Planned Parenthood Federation of AmericaAriel Kaplan / Planned ParenthoodView session - Breakout session
What changes when 300 dbt users at Virgin Media O2 move to Fusion
Melissa Simpson / Virgin Media O2Jason Jones / Virgin Media O2View session
