Don't ship what you can't measure: Evaluation-driven development for AI analytics agents
Building AI analytics agents is becoming democratised; tools like Snowflake Cortex Analyst make it easier than ever to deploy a text-to-SQL agent over your data models. However, trusting the output is still a challenge, and evaluating AI agents in analytics remains a bottleneck for most teams. As a rule of thumb: don't ship what you can't measure.
This talk introduces evaluation-driven development for AI analytics agents: define what "correct" means before you build, measure every iteration, and only deploy when the evidence proves it's ready. Built from deploying Cortex Analyst agents leveraging modelled data in dbt across healthcare and e-commerce clients, this methodology is practical and repeatable.
You'll learn how to design golden question sets that test real business scenarios, build automated eval pipelines scoring across six dimensions, and track accuracy improvement across iterative semantic view enrichment.
Check out more sessions
- View sessionNetworking
Data leaders behind the bar: Learn to create cocktails with Analytics8 & ThoughtSpot
- Training add-on
{SOLD OUT} Governed & scalable AI-assisted analytics with dbt
Jessica Stayton / dbt LabsRaini Laughlin / dbt LabsView session - Lightning talk
The limits of DIY: how an S&P 500 Company escaped the data context trap
Duncan Gilchrist / DelphinaMark Goodwin / Gen DigitalView session
