September 15-18
The Cosmopolitan
Las Vegas

Don't ship what you can't measure: Evaluation-driven development for AI analytics agents

Friday, September 1812:00 PM PT

Building AI analytics agents is becoming democratised; tools like Snowflake Cortex Analyst make it easier than ever to deploy a text-to-SQL agent over your data models. However, trusting the output is still a challenge, and evaluating AI agents in analytics remains a bottleneck for most teams. As a rule of thumb: don't ship what you can't measure.This talk introduces evaluation-driven development for AI analytics agents: define what "correct" means before you build, measure every iteration, and only deploy when the evidence proves it's ready. Built from deploying Cortex Analyst agents leveraging modelled data in dbt across healthcare and e-commerce clients, this methodology is practical and repeatable.

You'll learn how to design golden question sets that test real business scenarios, build automated eval pipelines scoring across six dimensions, and track accuracy improvement across iterative semantic view enrichment.

Check out more sessions

View all sessions
dbt Summit on stage

Ready to join us?

Join data leaders and practitioners at dbt Summit for three days of ideas, skills, and shared progress.