# Section: Blog — 50 most recent posts --- title: "Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026" description: "dbt v2, dbt State are GA and Fivetran + dbt Labs debuts Fivetran Context Layer, dbt Charts and a new open lakehouse vision" url: "https://www.getdbt.com/blog/fivetran-dbt-labs-announces-new-capabilities-to-make-enterprise-data-agent-ready-at-dbt-summit" date: "2026-09-16" authors: ["Elaine Green"] categories: ["Press"] --- # Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026 _Fivetran + dbt Labs makes dbt v2 and dbt State generally available, alongside the debut of Fivetran Context Layer, dbt Charts and a new open lakehouse vision for greater flexibility across storage and compute_ **LAS VEGAS – September 16, 2026** – Fivetran + dbt Labs today announced the general availability of dbt v2 and dbt State, delivering new levels of speed and cost optimization, and introduced Fivetran Context Layer, new dbt Wizard experiences, dbt Charts and an open lakehouse vision for greater flexibility across storage and compute. As enterprises deploy AI agents into production, they need trusted data, context the business controls, and the flexibility to work across the platforms, models and tools already in use. To meet these demands, Fivetran + dbt Labs is advancing its vision for Open Data Infrastructure: a vendor-neutral, interoperable architecture that lets organizations independently choose and evolve their storage, compute, data movement, transformation and visualization technologies at every layer. This enables AI systems to work across platforms using data and context the enterprise owns, not a single vendor. "Every model our customers have built, every test they've written, every metric they've defined already captures the context AI agents need to do meaningful work," said Anjan Kundavaram, Chief Product Officer, Fivetran + dbt Labs. "What we're delivering now is the open infrastructure to put that context to work across systems, while giving organizations the freedom to choose how their data is stored, moved, transformed and used as AI evolves." **A faster, more efficient engine with deeper SQL understanding** [dbt v2](https://docs.getdbt.com/blog/dbt-v2-is-ga) is a full Rust rewrite of the dbt engine built for the scale that teams run at today and for how agents write SQL. It parses a 10,000-model project up to 10x faster than v1 and gives teams and their agents accurate real-time feedback, surfacing errors, column checks, and lineage before anything runs. With this release, the two engine era of Core and Fusion ends. Now, dbt is one engine with two versions: dbt Core v1, the python implementation, is dbt v1. Fusion, the Rust implementation, has become dbt v2. Both versions remain Apache 2.0-licensed and security-supported. [dbt State](https://www.getdbt.com/blog/dbt-state-is-ga) determines what has changed by checking warehouse metadata and model SQL, then builds, skips, clones or defers each run accordingly. This simplifies orchestration and allows engineers to iterate faster without complex development rituals, while reducing unnecessary warehouse compute. "dbt State has been a paradigm shift for how we work,” said Gordon Curzon, Head of Analytics Engineering, Virgin Media O2. “With freshness codified, simpler orchestration, and freed-up developer capacity, we focus more time on initiatives that add value to our business on top of the 25% savings on both job run time and BigQuery compute costs." **More flexibility across storage and compute** Open Data Infrastructure centers on a customer-owned data layer built on open formats, avoiding vendor lock-in for storage and compute so organizations can store data once and access it through different engines for different use cases. Fivetran's Managed Data Lake Service, already generally available, organizes, structures and maintains data as managed Apache Iceberg™ tables in customers' own cloud storage. Lake Compute, now in Private Beta, is a single-node SQL engine built on DuckDB and runs dbt models directly against Apache Iceberg™ tables, built and priced specifically for transformation, not general-purpose compute. Together, the two give teams the flexibility to run each workload on whichever engine fits best, optimizing cost without re-platforming. **An open standard for agent context** Agents are only as trustworthy as the context they can access. Fivetran Context Layer (Private Beta) unifies the data and metadata needed to give LLMs and AI agents relevant context, building on dbt’s structured context and adding unstructured knowledge, like docs and Slack threads. This service uses Agents Schema, an open source standard, that centralizes context in a structured, extensible format directly in the data warehouse. The context is accessible to teams via preferred MCP or AI tools, including generally available integrations through AI marketplaces including Anthropic and a plugin in ChatGPT. **One agent, grounded in your dbt project – wherever you work** Coding agents can now write SQL as well as most engineers, but writing code isn't the same as understanding a governed dbt project, including its lineage, its tests, its contracts, and what breaks when something changes. dbt Wizard in the dbt platform (Public Preview) is built to close that gap. It's natively connected to your project, knows which tool to call, pulls the right context automatically, and proactively validates changes before they ship. Wizard is also expanding beyond the dbt platform with Wizard CLI (Public Beta), bringing the project-grounded agent directly into the terminal, and Wizard Desktop (Private Beta), a dedicated local workspace for longer, more complex work. Wizard Explore Mode (Public Preview) brings conversational analytics to business users, enabling them to ask questions in plain language and get answers grounded in the same dbt project the data team maintains. When an answer falls short, those questions can also surface what the data team should improve next. **A shared language for BI, built for humans and agents** [dbt Charts](https://docs.dbtcharts.com/) (Public Beta) brings governed BI alongside the models it depends on. Instead of governance living in a separate, closed tool, they are defined as YAML and version-controlled alongside the dbt models they reference, creating a shared, declarative format that both humans and AI agents can read, write and review. **Customers building with Fivetran + dbt Labs** “Since rolling out dbt State, we’ve reduced warehouse costs by 59% on scheduled jobs in dbt platform,” said Chris Shepherd, Principal Data Engineer, RxBenefits. “That’s $8,173.23 in the first 60 days alone on top of a Snowflake adaptive warehouse. We’ve reused 716k models instead of rebuilding, which reduced query run time a total of 14 days, 11 hours, and 15 minutes over the same period.” “We've been impressed by the flexibility Lake Compute gives us. Now we can choose where each dbt workload runs, and use whichever engine actually fits the job,” said Tyson Doberneck, Senior Data Engineer, Obie. "dbt Wizard is changing how we work. Instead of hand-coding everything, we draft logic with an agent that already has full context on our jobs and our codebase. No more finding and uploading a manifest file just to explain myself, that step used to slow down every request. Now we're pointing it at sales and marketing data too, so analysts get answers themselves instead of waiting on my team," said Farin Fukunaga, Data Engineering Lead, Paylocity. To learn more about the product innovations unveiled at dbt Summit 2026, read the [recap](https://www.getdbt.com/blog/dbt-summit-2026-product-announcements) or register to watch the full keynote: [https://www.getdbt.com/dbt-summit/registration/online](https://www.getdbt.com/dbt-summit/registration/online) **About Fivetran + dbt Labs** Fivetran + dbt Labs deliver the data infrastructure layer that makes agents trustworthy – from the moment data moves, through every transformation, to the context an agent reasons from. The Fivetran platform moves, manages, and transforms data from every system a business runs on into a secure, reliable foundation engineered to evolve, with the flexibility to work across clouds, engines, and tools. With Fivetran, analytics, operations, and AI run on data you trust and control. Thousands of organizations worldwide, including OpenAI, LVMH, Pfizer, and Verizon, rely on Fivetran to turn data into a competitive advantage. Learn more at Fivetran.com, or follow Fivetran on LinkedIn. Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Globally, more than 100,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on LinkedIn, X, Instagram, and YouTube. --- --- title: "Everything we announced at dbt Summit and why it matters" description: "Every product announced at dbt Summit, from dbt v2 and dbt State to dbt Wizard and dbt Charts, and why each one matters." url: "https://www.getdbt.com/blog/dbt-summit-2026-product-announcements" date: "2026-09-16" authors: ["Corinne Hallander"] categories: ["Product"] --- # Everything we announced at dbt Summit and why it matters A decade ago, dbt gave people a name for something they were already trying to do: write SQL like software engineers, version it, test it, document it, trust it. That idea became a new practice: analytics engineering. The people who built their careers on it became some of the most capable data professionals in the industry. Now, those same people are asking a harder question: what impact will AI have on the practice of analytics engineering? AI needs an engine that's fast enough to keep up, context that's‌ trustworthy, and agents that understand your business instead of guessing at it. The question isn't whether analytics engineers still matter; it’s how do they level up for this new era? This week at [dbt Summit](https://www.getdbt.com/dbt-summit/registration/online), we welcomed 2,000 data professionals to Las Vegas (and thousands more online) to discuss the future of analytics in the AI age. Fivetran and dbt Labs are coming together around one thesis: the data foundation that makes analytics trustworthy is the same foundation that makes AI trustworthy. To help our users level up on both fronts, we announced a series of new features across the dbt and Fivetran product portfolios. ## Level up the engine ### One dbt, one engine, built to move as fast as you do dbt has always evolved with growing data workloads. Last year, we introduced the dbt Fusion engine, a full rewrite of dbt in Rust with native SQL comprehension and dramatically faster performance than the original Python-based standard. But maintaining two engines created real friction, both for us and for the thousands of teams trying to figure out which one to build on. **dbt v2** ends that. Now GA, [dbt v2](https://docs.getdbt.com/blog/dbt-v2-is-ga) is one modern, Rust-based engine powering all of dbt, whether you work locally or in the dbt platform. It's the same workflow, on a faster foundation—thanks to all the innovation in Fusion over the past 18 months. ```json { "_key": "c077b514b103", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/_6ZMtNWheiQ" } ``` We publish [two distributions of v2](https://docs.getdbt.com/blog/comparing-dbt-and-dbt-oss): - The superset `dbt `includes all parts of the framework, super-charged with SQL comprehension features. It’s free to use, with additional optional paid features (for example, dbt State). - The subset distribution `dbt-oss` includes _only_ the components that have an Apache 2 license dbt Core isn't going anywhere either. It’s still Apache 2.0, still open source, simply renamed as the previous version, dbt v1. Adapter availability: - BigQuery, Databricks, DuckDB, Redshift, and Snowflake are GA - ClickHouse and Spark are in beta and available to install and test locally - Athena, Fabric, and Postgres are coming soon > “With the new dbt v2 engine, the performance improvements showed up across the entire development experience. Teams spent less time waiting for processes to complete, moved changes through the pipeline faster, and could focus more of their time on building and delivering data products.” — Vishesh Jain, Delivery Lead, Data Analytics Platform at RMIT University [Get started](https://docs.getdbt.com/docs/local/install-dbt?install-method=pip&version=2) on v2 today. ## Build what's changed, skip what hasn't with dbt State, now GA **Now officially GA**,** [dbt State](https://www.getdbt.com/blog/dbt-state-is-ga)** checks your warehouse metadata and model SQL for what's changed, then builds, skips, clones, or defers each run accordingly. That intelligence results in an average 15-30%+ reduction in warehouse compute and removes the manual syntax and workarounds teams built to avoid running more than they needed to. > “Since rolling out dbt State, we’ve reduced warehouse costs by 59% on scheduled jobs in the dbt platform, which runs on top of a Snowflake adaptive warehouse. We’ve reused over 700k models instead of rebuilding, which reduced query run time by two weeks over a 60 day period.” — Chris Shepherd, Principal Data Engineer, RxBenefits With dbt State, freshness moves from the job to the model. Every model carries its own freshness requirement in code, a `lag_tolerance` that says how stale it's allowed to be, simplifying orchestration. And because codified freshness rules and automatic reuse constrain what any run can cost, the guardrails now sit in the infrastructure instead of with whoever, or whichever agent, issues the command. The result is felt daily in development. dbt State removes complex dev setup rituals and risks of accidental builds for faster, safer dev work. > “dbt State has been a paradigm shift for how we work. With freshness codified, simpler orchestration, and ultimately, freed-up developer capacity, we focus more time on initiatives that add value to our business. And that’s on top of the 25% savings on both job run time and BigQuery compute costs.” — Gordon Curzon, Head of Analytics Engineering, Virgin Media O2 ```json { "_key": "eb8188a221e6", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/mczQsM8pl78" } ``` dbt State runs on dbt v1.7 through v2, wherever you run dbt: locally, in the dbt platform, with your own orchestrator, and across Snowflake, BigQuery, Databricks, and Redshift. [Get started on dbt State today](https://docs.getdbt.com/docs/deploy/dbt-state-setup?version=2) so you can optimize costs, save time, and level up wherever you run dbt. ## Own your data. Align costs to value. Introducing Lake Compute, now in Beta Most companies run every dbt model, from massive joins to simple staging tables, on the same warehouse compute. With the rise of Apache Iceberg, you can store your data in open-table formats, and have the flexibility to choose the right compute engine for each workload. We now offer that choice with **Lake Compute (Private Beta)**,** **a single-node SQL engine built on DuckDB for dbt that runs transformations directly against Apache Iceberg tables. Tag one model or a hundred to run on Lake Compute. The rest keep running on your warehouse, and `ref`s keep working across both. Engine choice becomes a per-model decision instead of a replatforming program. With Lake Compute, we’re working towards the promise of an open data lakehouse: one where you can choose the right compute for the job. Learn more about Lake Compute [here](https://docs.getdbt.com/docs/lake-compute). Want to see it in action? Join our [upcoming webinar](https://www.getdbt.com/resources/webinars/cost-optimized-transformations-with-dbt-and-apache-iceberg-on-open-multi-engine-compute) on cost-optimized transformations with dbt, Apache Iceberg, and multi-engine compute. We'll walk through a live, dual-engine dbt project running partial Snowflake, partial Lake Compute, on the same Iceberg tables, so you can see exactly what "engine choice as a per-model decision" looks like in practice. [Save your seat here](https://www.getdbt.com/resources/webinars/cost-optimized-transformations-with-dbt-and-apache-iceberg-on-open-multi-engine-compute). ## Level up for AI: context, engineered AI agents don't necessarily need _more_ of your data. They need an understanding of what the data means, where it came from, whether it's fresh, and who owns it. That foundation is already in place: [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl?version=2) governs metrics, [Agents Schema](https://www.fivetran.com/blog/how-agents-schema-brings-trusted-business-context-to-ai) centralizes context, and [dbt MCP Server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2) exposes your models, metrics, lineage, and test results to any AI agent. Delivering structured dbt context to your favorite AI tools just got easier with new **out-of-the-box integrations with Anthropic and a plugin in ChatGPT (GA)**. No more multiple MCP servers to manage. Just one click, and your team can securely access structured context from your dbt project instantly. Structured context gets you closer to reliable AI answers. But most of what a business runs on doesn't live in structured tables at all. It lives in call recordings, support tickets, Slack threads, and it changes constantly. That's what [**Fivetran Context Layer**](https://fivetran.com/docs/context-layer) is built for: turning every source your business runs on—both structured sources like dbt and unstructured ones that BI tools never touched—into context for the AI tools you're already using, from Claude to Slack. Because Fivetran and dbt have visibility into your entire data estate, this context gets built and maintained as data moves and transforms, not stitched together after the fact. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0864156bd3668028d1019a580b859b1e953f9568-2048x1342.jpg) [Sign up](https://go.fivetran.com/signup/context-layer) for the early access program. Want to go deeper on this? Join our session, [_From Analytics Engineer to Context Engineer: A dbt Playbook_](https://www.getdbt.com/resources/webinars/from-analytics-engineer-to-context-engineer-a-dbt-playbook), where we make the case that context engineering is analytics engineering with a new last mile. We'll walk through the exact patterns dbt Labs uses internally to turn unstructured sources into governed, versioned context, and introduce the `dbt_context_engineering` package that puts those patterns into practice. [Save your seat here](https://www.getdbt.com/resources/webinars/from-analytics-engineer-to-context-engineer-a-dbt-playbook). ## Level up with AI AI is raising the bar on what it takes to ship trustworthy data: more models, more context, less room to guess. Most of that time gets lost relearning what should already be known: agents rediscovering a project's lineage and tests before they can start, visualizations living outside the codebase entirely. [**dbt Wizard**](https://www.getdbt.com/product/dbt-wizard) is an agent built specifically for analytics engineering. It’s grounded natively in your project, so it already knows which tools to call and which context to pull without any setup. It validates proactively: checking upstream and downstream impact, compiling and building the change before anyone sees the diff. And because data work is visual, you can review all of that against the full DAG. Wizard in the dbt platform is now in Public Preview. [**dbt Wizard Explore Mode**](https://docs.getdbt.com/docs/platform/wizard-home#ask-questions-in-explore-mode) (Public Preview) brings conversational analytics right where the data work already happens. It lets business users and data teams ask questions in plain language and get answers grounded in the same dbt project your data team already maintains. When an answer falls short, you're already one step from the model that needs fixing. [**Wizard CLI**](https://docs.getdbt.com/docs/dbt-ai/wizard-cli) (Public Beta) puts the same project-grounded agent in the terminal you're already running dbt from. No new app, no new tab, and no platform account required. And [**Wizard Desktop**](https://docs.getdbt.com/docs/dbt-ai/wizard-desktop?version=2) (Private Beta) picks up where the terminal runs out of room: a dedicated local workspace for longer, more complex work. You can run several tasks side by side instead of juggling windows with a visual preview of the code, data, and lineage before anything ships. Learn more [here](https://www.getdbt.com/product/dbt-wizard). ```json { "_key": "0f25474437ae", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/cEwRVyYZd1I" } ``` ## Closing the last gap: dashboards as code Asking questions is only one way people need to work with data. Sometimes what they need is a dashboard—one they'll come back to every day. And that's where the story gets uncomfortable: everywhere else in the stack, teams have brought in real engineering discipline—version control, code review, CI. dbt was the reason SQL went from copy-pasted queries to models you can trust. But the dashboard at the end of that pipeline—the thing an executive looks at—was still locked in, trapped inside whatever proprietary tool you bought, stored in someone else's format, disconnected from the code that produces the numbers. [dbt Charts](https://docs.dbtcharts.com/) (Public Beta) closes that gap. It brings discipline to dashboards: built as YAML, version-controlled right next to the models they depend on, living in the same repo, the same pull request, the same CI as the SQL underneath it. Because it's declarative, it becomes a shared language, one humans and agents can both read, write, and review, just as reliably as they do your SQL. [Check out dbt Charts](https://dbtcharts.com/) today. ## We’re all leveling up These announcements represent more than a set of new products. Together, they represent the forward direction of Fivetran + dbt Labs. AI is changing the game, and we are building the data foundation to help you succeed in the agentic AI era. Here’s a recap of how to get started with any of these new products: - [Upgrade to dbt v2](https://docs.getdbt.com/docs/local/install-dbt?install-method=pip&version=2) for faster development, SQL comprehension, and more. - Turn on [dbt State ](https://www.getdbt.com/product/dbt-state)and start simplifying your development and saving on compute. - Use [Lake Compute](https://docs.getdbt.com/docs/lake-compute) to run the right engine for each model, without a replatforming project. - Start developing agentically with [dbt Wizard](https://www.getdbt.com/product/dbt-wizard)—it’s available in the [dbt platform](https://docs.getdbt.com/docs/dbt-ai/wizard-ide?version=2), in a brand new [Desktop app](https://docs.getdbt.com/docs/dbt-ai/wizard-desktop?version=2), and via the [CLI](https://docs.getdbt.com/docs/dbt-ai/wizard-cli). - [Explore mode](https://docs.getdbt.com/docs/platform/wizard-home#ask-questions-in-explore-mode) in Wizard makes it easy for business users to chat with your data. [Add a read-only user](https://docs.getdbt.com/docs/platform/wizard-read-only-users?version=2#set-your-team-up-for-good-answers) at no cost, they'll learn about the data, and you'll learn a lot from what they ask. - Get on the waitlist for [Fivetran Context Layer](https://go.fivetran.com/signup/context-layer). - Try [dbt Charts](https://dbtcharts.com/), the first language-first BI layer. --- --- title: "We built dbt State to stop rebuilding what hadn't changed" description: "dbt State is now generally available everywhere you run dbt. See how it cuts compute costs and speeds up development." url: "https://www.getdbt.com/blog/dbt-state-is-ga" date: "2026-09-16" authors: ["David Macias"] categories: ["Product"] --- # We built dbt State to stop rebuilding what hadn't changed ```json { "_key": "191612bc0735", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/mczQsM8pl78" } ``` Today [dbt State](https://www.getdbt.com/product/dbt-state) is generally available, everywhere you run dbt. This includes your own orchestrator: Airflow, Dagster, GitHub Actions, or a laptop. And as of today, it’s available on Snowflake’s dbt projects. It works on Snowflake, BigQuery, Databricks, and Redshift. Most dbt jobs run on a schedule. dbt State makes them run on change instead. On every run, it reads your model SQL and your warehouse metadata, works out whether a model's result would be different, and then builds, skips, clones, or defers each node accordingly. There's no selection syntax to write, no manifest scripts, and no custom orchestration to maintain We built it on a straightforward idea: an enormous amount of redundant work happens in data pipelines, and reusing what hasn't changed would save teams hours of run time and warehouse compute. That has held true. > “Since rolling out dbt State, we’ve reduced costs by 59% on scheduled jobs in the dbt platform running on a Snowflake adaptive warehouse. We’ve reused over 700k models instead of rebuilding, which reduced query run time by two weeks over a 60 day period.” — Chris Shepherd, Principal Data Engineer, RxBenefits On average, more than 22 million models are built on dbt every single day. Most of the time the upstream data hasn't refreshed, so thousands of models rebuild to produce a result identical to the run before. Early adopters have averaged 15-30% compute savings by stopping that. Simply turning it on nets some benefits, but the bigger jump in compute efficiency happens when teams apply tuned configurations, or move to an adaptive warehouse on Snowflake. More on that below. ## The savings are real, and they're not what teams are talking about Compute savings are the easiest thing to measure, which is why they're usually the reason you turn dbt State on. They show up on the next bill, and they're what gets the project approved. But when we went back to the teams who have been running dbt State in production _and development_ for weeks and asked what changed, almost nobody led with the number. > "dbt State created a paradigm shift in how we work. With capacity freed up, freshness codified, and simpler, yet smarter orchestration, we can focus on initiatives that add value to our business. And that's on top of the 25% savings on both job run time and BigQuery compute costs." — Gordon Curzon, Head of Analytics Engineering and Data Modeling, Virgin Media O2 ## Orchestration stops being a scheduling problem Before dbt State, the way to hit a freshness SLA is to build a job for it. Then another job, on a different schedule, for the part of the DAG with different requirements. Then tags to trigger the right subsets, selectors to narrow them, and source freshness gating layered on top to stop jobs firing when nothing has landed. It works, but it also becomes brittle over time. dbt State changes moves the freshness from the job schedule to the model. Instead of a job deciding what runs, every model carries its own freshness requirement in code, a `lag_tolerance` that says how stale a model is allowed to be. Then dbt State decides, per model, per run, whether that requirement is met. Instead of asking which job a model belongs to, you ask how fresh it needs to be. That's a question an analytics engineer can answer, in a pull request, with review. At Custom Ink, a four-person analytics engineering team supports around 25 analysts, data scientists, and machine learning engineers across a project of more than 1,000 models. Consolidating that into a single deployment is the part they keep coming back to. > "We've been able to simplify our previous setup down to one run, which is amazing, because it's a very clean deployment. Everybody knows what's going to run. There's no confusion." — Brett Petersen, Director of Data and Metrics, Custom Ink ## Development is where you notice it every day The second thing we heard is that dbt State changed development just as much as it changed deployment. In dev, the old ritual is familiar to anyone who has worked in a large project. You want to edit one model in the middle of a long lineage. So you clone production into your dev schema, or you write a macro to do it, or you remember to turn on deferral, or you skip all of that, rebuild the upstream chain, and wait. As best as you can, you track how stale your copy has become. dbt State does the cloning and reuse automatically. You build what you're working on, and the rest is reused. > "Now you just dbt build --select what you're working on. dbt compares against prod state, reuses every fresh upstream at zero compute, and rebuilds only your change. The whole chain resolves in seconds. No clone command, no waiting, no dev-environment ritual, the state tracking you used to do in your head is now the tool's job. It saves so much development time." — Mykkel Ryda**l**, Lead Data Engineer, Joe & The Juice At Joe & The Juice, iteration on a core sales fact table of 400M+ rows used to take 15-25 minutes a cycle, and `--full-refresh` was blocked in code because it was too expensive to allow. Setup before new dev work went from 15-25 minutes to zero. ## The guardrails move into the infrastructure For years, the safety net in dev has been knowledge. Nothing technically stops a dbt build from kicking off an entire pipeline. What stops it is a developer knowing the right selection syntax, which means the guardrail is only as good as the least experienced person running the command. It becomes untenable and harder to scale as development work becomes evermore agentic. dbt State is a structural guardrail rather than a manually enforced one. Codified freshness rules and automatic reuse constrain what any run can cost, regardless of who or which AI agent issued it. > "With dbt State doing the cloning and reuse automatically, the guardrails are just in the infrastructure now, whether it's an analyst or an agent doing the work." — Brett Petersen, Director of Data and Metrics, Custom Ink Custom Ink leans heavily on AI-assisted development, as many teams do these days, and one clean job flow is part of what makes that workable. It’s consistent DAG behavior at scale. ## What dbt State GA means dbt State is out of preview and generally available. Setup is turning it on. Run a job with dbt State enabled twice and see `No-Op` (no operation) show up over and over in your logs. Declare freshness SLAs or other configurations with precise controls: `lag_tolerance`, `require_fresh_data_from`, `compare_unrendered_code`, and more. Use `dbt state explain` to see why any node was built, skipped, cloned, or deferred. Make sure the rest of the data team turns on dbt State in local development so they can iterate faster and lower the cognitive overhead it takes to minimize risk of costly or potentially breaking builds. We worked with Snowflake to bring dbt State’s intelligent reuse to Snowflake dbt Projects, so as of today, those teams can skip unchanged models without moving where they work. dbt State on platform includes Cost Insights and now has a richer `explain` experience for more insight into how to optimize and prove ROI. It also manages concurrent builds so two jobs don't clash on the same model at the same time. Pricing is consumption-based and tied to reuse rather than builds. You pay for daily active target tables (DATT), so each distinct model or test that dbt State skips, clones, or reuses on a given day. Every reuse after the first within the same day is free. Price is independent of table size and compute, so running jobs more often doesn't increase what you pay. In fact, running more often tends to save more. And every user gets a 30-day free trial. Custom pricing is available for upfront commits. See the [docs](https://docs.getdbt.com/docs/deploy/dbt-state-setup?version=2) to get started today. --- --- title: "Celebrating the 2026 dbt partner of the year winners" description: "Meet the 2026 dbt Labs Partner of the Year winners: phData, Snowflake, Cívica, Datum Studio, and 66degrees." url: "https://www.getdbt.com/blog/2026-partner-of-the-year-winners" date: "2026-09-15" authors: ["Yahsmene Butler"] categories: ["Partnerships"] --- # Celebrating the 2026 dbt partner of the year winners Our partners help data teams level up their dbt implementations: faster rollouts, cleaner data foundations, and AI initiatives that ship. At Partner Day during dbt Summit 2026, we recognized five partners who did that work at the highest level this year. Here’s who won, and why. ## Partner of the year: phData phData earned partner of the year for the fourth consecutive year for what its customers walk away with: dbt implementations and data foundations to build trusted agents on, and a certified bench ready to take on the hardest problems. As a Visionary-tier partner, phData is a team that shared customers trust with their most critical data programs. > “Winning dbt Labs Partner of the Year for the fourth year in a row says more about the depth of this partnership than any single project could. dbt has become the backbone of how we help clients turn raw data into something the business can actually trust and reason with. We’re proud of the work we’ve done together, and even more excited about where dbt is headed as it becomes central to how our clients build agentic and AI use cases on top of their data.” — Dustin Dorsey, Senior Director, Data Engineering, phData ## Technology partner of the year: Snowflake Snowflake earned technology partner of the year for the joint value that customers get out of the partnership. In every region, teams running dbt on Snowflake are standing up trusted, production-grade pipelines faster, with dbt Labs and Snowflake engineers working the same problems together rather than handing customers off to each other. It’s one of the most widely adopted foundations in the ecosystem. > “Snowflake and dbt Labs share a vision of empowering data teams to move faster with confidence, and this partnership continues to raise the bar for what our joint customers can achieve. Together, we’ve made it simpler than ever for organizations to build trusted, production-grade data pipelines on Snowflake. We’re proud to be named dbt Labs Partner of the Year and excited to deepen our collaboration to help even more customers turn their data into real business impact.” — Rodrigo Rocha, VP, Global ISV and Technology Partnerships, Snowflake ## EMEA partner of the year: Cívica Cívica earned EMEA partner of the year for the quality of what customers get, not just the volume. More customers in the region do their dbt work with Cívica than with any other partner, and every engagement is backed by a certified delivery team. It’s a Visionary-tier partner customers across EMEA rely on to get it right the first time. > “Being recognized as dbt Labs Partner of the Year in EMEA is a tremendous honor and a true testament to the hard work and dedication of our entire team. What excites us most about collaborating with dbt is the strength of its community: a vibrant ecosystem where open collaboration, innovation, and shared learning help drive the entire industry forward. We’re more motivated than ever to continue accelerating the future of data alongside this incredible network.” — Josep Roig, CEO, Cívica ## APJ partner of the year: Datum Studio Datum Studio earned APJ partner of the year for the depth of what it delivered for customers in Japan this year: more joint dbt work than any other partner in the region, as many customer wins as any partner globally, and one of the strongest certified delivery teams in APJ. In a market where technical depth is everything, Datum Studio was the team customers turned to this year. > “dbt has become the de facto standard for analytics engineering, and our role is to make that standard take root in the Japanese market, not just as a tool, but as a way of working. What excites me most about the partnership is where it’s heading: as AI agents start writing and reviewing transformations, the foundation dbt provides, lineage, testing, documentation, governance, and shared context, is what makes that work trustworthy at scale.” — Sohei Takechi, President and CEO, Datum Studio ## Emerging partner of the year: 66degrees 66degrees earned emerging partner of the year for how fast it got to real customer work. In under a year, 66degrees went from a standing start to delivering dbt projects for customers, bringing Google Cloud depth to shared accounts, and showing up alongside dbt Labs at field and virtual events. This is what a partnership looks like when a team decides to invest. > “66degrees is honored by this recognition as dbt Labs’ Emerging Partner of the Year, which celebrates our long-term Fivetran partnership and our years spent building dbt directly into our core accelerators, including Hub66 and Paradigm Data, our agentic data product delivery agents. Fivetran and dbt together define a transformative direction for data architecture, establishing the seamless pipeline required for modern analytics and AI. As both platforms innovate, we look forward to evolving our accelerators and driving the next era of intelligent data delivery right alongside them, giving our customers a best-in-class data platform.” — Daniel Zagales, SVP, Data and Analytics, 66degrees ## What this means for the dbt community These five partners represent five different ways of showing up for customers: depth of delivery, regional expertise, and the speed to build something new. Congratulations to phData, Snowflake, Cívica, Datum Studio, and 66degrees on the well-earned recognition. [Become a dbt Labs partner today.](https://www.getdbt.com/partners) --- --- title: "Why your AI pilot stalled at the context gap" description: "Many AI projects are stuck. Don’t blame the models. The problem is a lack of trusted context. Here’s how you solve it." url: "https://www.getdbt.com/blog/why-your-ai-pilot-stalled-at-the-context-gap" date: "2026-08-27" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Why your AI pilot stalled at the context gap Everyone has high hopes for the value they can derive from their agentic AI projects. The dream is to get them into production where they can assist users and make autonomous decisions that drive the business forward. But many projects get stuck in the pilot phase. The numbers are stark: - [Only 16% of companies have deployed agentic AI](https://www.infosys.com/newsroom/features/2026/enterprises-scaled-agentic-ai.html) at an enterprise scale. - [Nearly 80% of companies in one survey](https://www.cloudera.com/about/news-and-blogs/press-releases/2026-04-14-nearly-80-percent-of-enterprises-say-ai-is-held-back-by-data-access-challenges-cloudera-report-finds.html) said their AI projects are constrained by data access issues - [70% of those asked by Deloitte](https://www.deloitte.com/us/en/about/press-room/deloitte-survey-examines-ai-readiness-agentic-ai-success.html) said they don’t feel they can adequately trust and govern agents This results in an all-too-familiar situation where AI agents either guess about data or make it up completely. An agent that finds multiple conflicting definitions of “revenue” across data sources might arbitrarily pick one. The problem occurs when AI agents are deprived of the governed context they need to make informed decisions. Let’s look at this context problem, why ungoverned agents fail, and how the dbt platform enables you to build scalable AI agents that everyone in your company can trust. ## The four ways an ungoverned agent fails Without trusted, governed data, [an agentic AI solution that works under test conditions often fails](https://www.getdbt.com/blog/why-agentics-projects-fail-and-how-to-fix-them) when faced with real user questions. There are four common reasons why: **It writes unreliable SQL**. This isn’t often an outright syntactic failure; usually, the SQL parses and runs. The problem is that it’s selecting the wrong values. **It invents or misreads metric definitions**. Without sufficient context, an agent might use stale, missing, or fabricated metrics. This problem is exacerbated by a lack of reviews. **It has no guardrails and no audit trail**. No auxiliary processes check the agent’s work, and there’s no record you can check to verify how it reached its conclusions. **It creates rising compute costs due to inefficient work**. Left to their own devices, AI agents may use more compute than necessary, causing your data processing costs to spike. They’re getting the work done, but processing is eating your profits. None of these are “sometimes agents make a mistake” issues. These are predictable and, fortunately, fixable problems. You just need to take the correct approach to data. For example: - SQL generation can be validated through rigorous testing. You can also train your models and agents to learn how your data works. - Revenue can be centrally defined and shared across teams using a [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction), instead of spread across dozens of data stores. ## Machine-readable governance In the past, we relied on humans to manually review data and ensure its accuracy. We’re producing too much data for that to be a scalable approach in the AI age. To make agentic AI truly scalable, you need governance. But not the type where everything is written down in a large document no one reads. You need **computational** and **machine-readable** governance. In a computational model, governance is enforced using several capabilities: - [Contracts](https://roundup.getdbt.com/p/contracts-have-consequences). A contract is a machine-readable description of how the data is shaped, how it functions, and how it differs between releases, along with the endpoints used to access it. Data that doesn’t meet a contract fails to ship, keeping a class of errors out of production. The contract also enables agents to discover and use the data, particularly as its shape evolves over time. - [Tests](https://www.getdbt.com/blog/data-testing). Data needs to be tested the same way we test software. Tests that ensure correct data can be run when shipping new data transformation changes, and run periodically in production to ensure ongoing data health. - [Semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction). A semantic layer provides one central location for all metric definitions. This eliminates the agent from guessing what “revenue” means. Each metric provides additional metadata that’s invaluable to AI agents: where it came from, how it was derived, and its business purpose. Without this machine-readable approach to governance, you can’t guarantee that your AI agents will return accurate answers at scale. ## The shift to agent consumption of data Historically, the primary consumers of data have been humans. It’s quickly becoming AI agents, which we humans now rely on to help distill the vast amounts of information we keep generating. [Fivetran and dbt Labs realized that, together, we could do more to advance a new era of trusted, Open Data Infrastructure for AI at scale](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents). Your AI agents don’t need to languish in the prototype phase. With the dbt platform, you have the tools you need to create high-quality, governed, and trusted data that enables your agents to make accurate decisions. To learn how, [watch the full demo of the dbt platform in action](https://www.getdbt.com/resources/webinars/dbt-platform-live-demo). --- --- title: "Scaling AI is easy. Trusting it is hard." description: "As AI scales, cracks in data trust and governance start to show. Here's what breaks first, and what it takes to fix it." url: "https://www.getdbt.com/blog/scaling-ai-is-easy-trusting-it-is-hard" date: "2026-08-25" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Scaling AI is easy. Trusting it is hard. Organizations are moving beyond experimentation and embedding AI into everyday business operations. As adoption accelerates, many are discovering that scaling AI successfully requires far more than deploying increasingly powerful models. Over the next three years, [92% of companies plan to increase their AI investments, yet only 1% consider themselves mature in AI deployment](https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/superagency-in-the-workplace-empowering-people-to-unlock-ais-full-potential-at-work). As organizations scale AI and agents, trusted data infrastructure becomes the foundation for trusted AI—from copilots to autonomous agents. Without trusted data, governance, and business context, even the most advanced AI systems struggle to deliver reliable, explainable outcomes. We explored these foundational principles in [The Data Leader's Primer for Agentic AI.](https://www.getdbt.com/resources/the-data-leader-s-primer-for-agentic-ai) Here, we examine what starts to break when those foundations aren't in place and why AI maturity has become the next challenge organizations need to solve. What starts to break as AI scales? As AI becomes embedded across more workflows and business functions, existing weaknesses become more visible and more pronounced, turning what were once manageable issues into enterprise-wide challenges. Organizations often experience familiar operational challenges such as: - Data quality issues become amplified as AI consumes more data and influences more decisions. - Data governance becomes harder to maintain across teams, systems, and AI workflows. - Ownership becomes unclear as responsibility for data, metrics, AI outputs, and agent workflows spans multiple stakeholders, particularly as agents begin operating autonomously across team boundaries. - Infrastructure, compute, and operational costs become harder to manage as AI adoption grows. - Trust becomes harder to maintain as AI-generated outputs and agent-driven actions reach more employees and customers. - Teams struggle to explain how AI-generated answers were produced. These issues rarely occur in isolation. Together, they point to the same underlying challenge: organizations are scaling AI faster than their trusted data infrastructure can support it. ## Why these challenges matter Whether organizations are using off-the-shelf AI tools, sophisticated agent harnesses, or advanced agentic workflows, AI depends on trusted data, governance, and business context to produce reliable outcomes. As AI adoption grows, the operational challenges that have long affected analytics become even more visible and more consequential. The further organizations move toward autonomous, agentic systems, the higher the cost of getting these foundations wrong. The findings from the [2026 dbt Labs State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) reinforce this reality: - **53%** report poor data quality as a top challenge. - **41%** cite ambiguous data ownership as an ongoing challenge. - **71%** are concerned about hallucinated or incorrect data reaching stakeholders. These findings reinforce that AI readiness depends on the trusted data infrastructure supporting it. As organizations scale AI and agents, the quality of that foundation increasingly determines whether AI can deliver reliable business outcomes at scale. ## Building the foundation for AI at scale Deploying more models is only part of what it takes to scale AI successfully. Organizations also need the trusted data infrastructure that allows AI to perform reliably over time. That foundation extends beyond data alone. It includes governance, ownership, business context, interoperability, and the operational efficiency required to help AI produce accurate, explainable, and consistent outcomes at scale. Together, these capabilities enable organizations to move beyond isolated AI initiatives and scale AI confidently across the business. Understanding your organization's current capabilities is the first step toward identifying operational gaps and prioritizing the investments that will have the greatest impact. Strengthening that foundation better positions organizations to support the next generation of AI systems and agents as technologies and use cases continue to evolve. ## What's next? The upcoming **Enterprise AI Data Maturity Model **provides a practical framework for assessing your organization's AI maturity, identifying capability gaps, and understanding where to focus next. The accompanying guide explores the capabilities organizations need to progress from trusted data to trusted AI and agents. Join dbt Labs Senior Director of Product Strategy Russell Christopher and Infinite Lambda Chief Product Officer Petyo Pahunchev on September 2 or 3 for a first look at the Enterprise AI Data Maturity Model—a five-stage framework for finding out where your organization stands and what it takes to move up. [Register for the webinar](https://www.getdbt.com/resources/webinars/from-data-chaos-to-agent-accessible-how-to-move-up-the-ai-data-maturity-curve). --- --- title: "Databricks processes your data. dbt defines what it means" description: "Your compute platform and your transformation logic are two separate decisions. Most executives approve them as one." url: "https://www.getdbt.com/blog/databricks-processes-your-data-dbt-defines-what-it-means" date: "2026-08-17" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Databricks processes your data. dbt defines what it means When Databricks announced it was acquiring Neon in May 2025, it published a number from Neon’s own telemetry: over 80% of the databases on the platform were being created by AI agents rather than by people. That number is about Neon. The question underneath it applies to every platform, Databricks included. When an agent writes your revenue model, someone still has to own the definition. When your CFO asks how last quarter’s number was produced, the answer has to live somewhere you can reach. This is for data leaders whose teams already run on Databricks and are being asked whether dbt is still worth the overhead. Short answer: Databricks is an excellent platform, and we’re partners. The dbt-databricks adapter is maintained by the Databricks team. If you’re running ML pipelines, building a lakehouse, or doing serious AI feature engineering, it belongs in your stack. Thousands of data teams run [dbt and Databricks together in production](https://www.getdbt.com/data-platforms/databricks). What I want to challenge is a different assumption: that choosing Databricks as your compute platform and choosing Databricks to own your transformation logic are the same decision. Most companies make both at once without noticing. They’re not the same decision. ## Databricks is good, and that’s not the point Lakehouse compute, Delta tables, collaborative notebooks, ML pipelines. These are real capabilities that deliver real value. Databricks has also been building products that sit squarely in the transformation layer: Lakeflow Declarative Pipelines (formerly Delta Live Tables) for pipeline orchestration, Databricks SQL for query and analytics work, Unity Catalog for governance and lineage. Each is reasonable on its own. Together they represent a set of decisions about where your transformation logic lives. Choosing Databricks for compute is an infrastructure decision. Choosing Databricks to own your transformation logic is a strategic bet on where your data team’s institutional knowledge will live. Most executives sign off on the first without realizing they’ve also made the second. ## What owning the transformation layer actually means The risks of full consolidation aren’t theoretical. They become concrete the first time you try to leave, audit, or explain. Start with portability. If your transformation logic lives in Lakeflow pipelines and Databricks notebooks, moving it means rewriting it. That exit cost won’t appear in today’s contract negotiation. It’ll appear as a multi-month engineering project you didn’t budget for, triggered by a pricing change or a strategic pivot that seemed distant when you signed. Then auditability. When your CFO asks why Q3 revenue was revised, tracing the answer requires Databricks platform access rather than a git diff. Version control of transformation logic depends on the platform, which means it depends on a vendor’s access controls and pricing tier. And timing. Enterprises don’t feel vendor lock-in when things are working well. They feel it when pricing changes, when they want to run workloads somewhere else, or when the platform pivots. By then the transformation layer is load-bearing infrastructure. Databricks would counter that Unity Catalog provides lineage and governance. That’s true. But lineage in a proprietary catalog is a record of what happened. Transformation logic in open, version-controlled SQL gives you the ability to change what happens, and to do it independently, anywhere. ## What dbt gives you that Databricks can’t replace dbt solves a different problem than Lakeflow Declarative Pipelines. It puts transformation logic in open, version-controlled SQL that any engineer can read, test, and run, regardless of what’s underneath. Five things dbt gives you that a Lakeflow pipeline can’t. **One project, any warehouse.** The same dbt project runs on Databricks, Snowflake, BigQuery, DuckDB, and a [growing list of compatible platforms](https://docs.getdbt.com/docs/trusted-adapters). The transformation logic is yours. When your warehouse strategy shifts, you port the logic instead of rewriting the pipelines. **Open standards, and the layer above them.** On June 1, 2026, dbt Labs open sourced the dbt Fusion engine runtime, releasing it as [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here) in alpha under the Apache 2.0 license. Databricks has made its own open-source moves here too, donating its declarative pipelines framework to Apache Spark in 2025. The difference is what sits on top. An open runtime is table stakes. What your team actually needs to own is the layer above it: the tests, the contracts, the metric definitions, and the lineage that explain what a number means and prove it hasn’t quietly changed. **A semantic layer that travels.** Metric definitions in the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) are platform-agnostic by design. Define revenue once and the definition holds whether the query lands on Databricks or somewhere else. **A community organized around a standard.** More than 100,000 data teams [build on dbt](https://www.getdbt.com/community). That community coheres around a shared, open standard for transformation rather than around a single vendor’s roadmap, and that’s worth weighing when you decide where your team’s institutional knowledge will live. **Cost control that isn’t tied to a platform.** [dbt State](https://www.getdbt.com/product/dbt-state) works as a caching layer for your pipelines: it builds what’s changed and skips what hasn’t. Teams see warehouse compute drop by 30% or more. It runs on dbt Core 1.7+, as a plugin or out of the box in the dbt platform, so the savings don’t depend on consolidating your transformation layer anywhere in particular. This is the shape of the argument in miniature. The efficiency comes from the engine understanding your project, not from the platform owning it. ## Five questions executives should ask before consolidating 1. If we needed to move 30% of our data workloads off Databricks in 90 days, what would that require? 2. Can we produce a complete audit trail for every transformation that produced last quarter’s revenue number, without Databricks access? 3. Are our metric definitions in code, or in Databricks notebooks? 4. Does our data team own the transformation logic, or does it live inside a vendor product? 5. If the pricing on our transformation tooling changes next year, what’s our fallback? Each of these describes a situation organizations have already faced at scale. The executives who ended up in difficult positions weren’t naive. They made the two decisions together, without realizing they were separate. ## The right boundary: what each layer owns “dbt vs. Databricks” is the wrong frame. We’re partners, and the combination is production-proven. The risk shows up when companies blur the boundary between the two. Here’s the boundary that holds. Layer What it owns Databricks Compute, storage via Delta, ML pipelines, AI feature engineering, collaborative notebooks dbt Transformation logic, the semantic layer, data contracts, lineage in code, testing Databricks processes your data. dbt defines what it means. Those are two different jobs, and conflating them is where the risk comes from. The executives who regret full consolidation are the ones who approved it while everything was working, before they needed to explain a number to the board, migrate a workload, or watch their data team’s institutional knowledge walk out the door embedded in Databricks notebooks. dbt is how your organization keeps the ability to reason about its own data, independently of any single vendor. **See how dbt and Databricks work together.** [Explore the integration](https://www.getdbt.com/data-platforms/databricks) --- --- title: "dbt Core v1.12 is GA" description: "dbt Core v1.12: what's new and how to upgrade." url: "https://www.getdbt.com/blog/dbt-core-v1-12-is-ga" date: "2026-08-17" authors: ["Grace Goheen", "Sara Gawlinski"] categories: ["Product"] --- # dbt Core v1.12 is GA dbt Core v1.12 is a big one. It does two things at once. It delivers meaningful improvements for teams using dbt Core today, including UDF enhancements, simpler Iceberg catalog and Semantic Layer specs, and plenty of quality-of-life upgrades like a new `on_error` config for handling upstream failures, a dedicated `vars.yml` file, and ad hoc SQL through `dbt run-operation --sql`. It also introduces a new opt-in Rust-based parser, the same one that powers [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0), giving teams a practical, low-risk way to start preparing for the next major version of dbt. Check out the [v1.12 upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2.0) for the full list of changes. If you’d rather hear it straight from the team, [join us for a live virtual recap and Q&A](https://www.getdbt.com/resources/webinars/dbt-core-v1-12-live). We’ll walk through what shipped, explain why it matters, share more context on the path to v2.0, and answer your questions live. Let’s get into what’s new in dbt Core. ### Take the first step toward dbt Core v2.0 with the v2 parser Before dbt can compile or run your project, it needs to read your project files, understand your resources and configurations, resolve dependencies, and construct the DAG. As projects grow, the time required to do that work can become a meaningful part of the development loop and startup time. dbt Core v1.12 introduces the opt-in `--use-v2-parser` flag, which delegates that work to the new Rust parser built for v2 instead of the Python parser used by dbt Core v1.x. On larger projects, the Rust parser can be 5–10× faster. The flag is entirely opt-in. Nothing changes unless you enable it, and you can return to the existing parser simply by removing the flag. That makes v1.12 a low-risk place to test the new parser against your real project, identify compatibility issues, and address them gradually rather than all at once. In other words: this is not just a performance improvement. It is a stepping stone to dbt Core v2.0. The [next major version of dbt Core](https://docs.getdbt.com/blog/dbt-core-v2-is-here) is being rebuilt in Rust on the same foundations as the dbt Fusion engine. It raises the baseline for parser performance, language validation, artifacts, documentation, and adapter development. Trying the parser in v1.12 gives you an early look at one of the most foundational parts of that new architecture without requiring you to move your whole project to dbt Core v2.0 today. And yes, we want your feedback. If you encounter a difference or edge case, open an issue and let us know. Real-world testing from the community is how we close those gaps before dbt Core v2.0 reaches GA. ## New dbt framework features ### More control when an upstream model fails This one has been a long time coming. In 2020, community member [@ian-whitestone proposed](https://github.com/dbt-labs/dbt-core/issues/2142) allowing downstream models to continue running in cases where an upstream failure did not make their results unusable. Let’s say for example an infrequently changing country or currency dimension: if that dimension fails to refresh, a daily rollup may still be able to process new transactions against its last successful version. The new [`on_error` config](https://docs.getdbt.com/reference/resource-configs/on_error?version=2.0) gives teams more control over whether downstream models should be skipped or allowed to continue after a failure. `on_error` accepts two values: - `skip_children` (the default): all downstream models are skipped, exactly as dbt behaves today. - `continue:` downstream models keep running instead of being skipped. ```sql -- models/dim_customers.sql {{ config( materialized='table', on_error='continue' ) }} ``` ### A cleaner home for project variables If your project contains a lot of variables, `dbt_project.yml` starts doing double duty: project configuration and variable storage, in one increasingly long file that everyone on the team edits. Back in 2020, [@benjaminsingleton suggested](https://github.com/dbt-labs/dbt-core/issues/2955) giving variables their own file to keep that file readable and cut down on merge conflicts. dbt Core v1.12 makes that possible with support for a dedicated `vars.yml` file at the project root. ```yaml # vars.yml vars: schema_name: analytics materialization: table ``` In addition to keeping `dbt_project.yml` cleaner, variables defined there are available while the project file is parsed, making them useful in project-level configuration as well. That means you can reference those variables inside your project file itself, and this is something you couldn't do when the variables lived in the same file they needed to configure: ```yaml # dbt_project.yml models: my_dbt_project: +schema: "{{ var('schema_name') }}" +materialized: "{{ var('materialization') }}" ``` ### Run ad hoc SQL without creating a macro The new `--sql` flag for `dbt run-operation` lets you execute a one-off database statement through dbt’s Jinja compilation context, without first creating a named macro. It is a simpler way to handle one-time operations while still using dbt’s existing connection and compilation behavior. ```sql dbt run-operation --sql "grant select on {{ ref('fct_orders') }} to role reporting" ``` ### Extend reusable logic with new UDF capabilities In dbt Core v1.11, user-defined functions (UDFs) officially became part of the dbt standard. That work was shaped by years of community experimentation and feedback and the community continued to help move it forward in v1.12. **JavaScript UDFs.** You can now define JavaScript UDFs for Snowflake and BigQuery directly within your dbt project. Drop the function body in a `.js` file under `functions/`: ``` // functions/is_positive_int.js return /^[0-9]+$/.test(a_string) ? 1 : 0; ``` Then define its arguments and return type in the corresponding properties file. A special thank you goes to @pempey, whose adapter override macro for experimenting with UDFs in additional languages gave the team a solid starting point for this work. **Python UDFs on Databricks.** Python UDFs aren't new, but starting in v1.12, they can run on Databricks (Unity Catalog required), joining Snowflake and BigQuery. **Multiple signatures with `overloads`.** The new `overloads` property lets one function accept several argument signatures, so you don't need a separate UDF for every input type. Each overload points to its own body file: ```yaml # functions/is_positive_int.yml functions: - name: is_positive_int arguments: - name: a_string data_type: string returns: data_type: integer overloads: - defined_in: is_positive_int_numeric arguments: - name: a_num data_type: numeric ``` **Third-party packages for Python UDFs.** Python UDFs can now declare public PyPI packages through the packages config. Your warehouse installs them when it creates the function: ```yaml # functions/is_positive_int.yml functions: - name: is_positive_int config: runtime_version: "3.11" entry_point: main packages: - numpy - pandas==1.5.0 ``` Together, these enhancements make reusable transformation logic easier to define, govern, and deploy alongside the rest of your dbt project. ### New Semantic Layer spec and Apache Ossie support v1.12 introduces the [latest dbt Semantic Layer YAML specification](https://docs.getdbt.com/docs/build/latest-metrics-spec?version=2), designed to make semantic definitions feel more closely connected to the models and columns they describe. Rather than defining a semantic model as a separate top-level resource, you can nest semantic information directly within a model. Entities and dimensions are defined at the column level, while simple metrics replace measures and can live alongside the model that provides their underlying data. Legacy spec: ```yaml semantic_models: - name: orders model: ref('orders') defaults: agg_time_dimension: ordered_at entities: - name: order type: primary expr: order_id - name: customer type: foreign expr: customer_id dimensions: - name: ordered_at type: time type_params: time_granularity: day - name: status type: categorical expr: order_status measures: - name: order_total agg: sum expr: amount metrics: - name: order_total type: simple type_params: measure: order_total ``` Latest spec: ```yaml models: - name: orders semantic_model: enabled: true agg_time_dimension: ordered_at columns: - name: order_id entity: type: primary name: order - name: customer_id entity: type: foreign name: customer - name: ordered_at granularity: day dimension: type: time - name: order_status dimension: type: categorical metrics: - name: order_total type: simple agg: sum expr: amount ``` This makes it easier to understand the relationship between a physical model and its semantic meaning without jumping between disconnected definitions. dbt Core v1.12 also adds support for defining semantic models using [Apache Ossie](https://www.getdbt.com/blog/osi-is-now-apache-ossie) documents (formerly the Open Semantic Interchange). dbt can parse Ossie-format JSON files alongside native dbt semantic models and generate an `osi_document.json` artifact representing your project’s Semantic Layer. Together, these changes move us forward on a path where semantic context is easier to author in dbt and more portable across the broader ecosystem. ## Adapter-specific features and enhancements As always, the Core release is only part of the story. Adapter maintainers have continued improving the experience across individual data platforms. Highlights include: - **Snowflake:** Iceberg v3 support and more control over dynamic-table scheduling, support for separate warehouses during initial builds, and transient tables. - **BigQuery:** Parallel microbatch execution, standard SQL for partition metadata, and per-resource execution timeouts. - **Redshift:** Support for the `query_group` session parameter, enabling better workload routing and query logging. - **Databricks:** Unity Catalog row filters and additive merging of tags across project configuration levels. Head to the [v1.12 upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2.0#adapter-specific-features-and-functionalities) for the complete adapter-by-adapter breakdown. ### Quick hits Here's what else landed in v1.12. - **Native private packages:** Install packages from private GitHub, GitLab, or Azure DevOps repositories using your existing SSH configuration, without specifying a full Git URL or separately configuring a token. - **Expanded Iceberg support:** Use a simplified `catalogs.yml` specification, enable cross-platform dbt Mesh, and create Iceberg v3 tables on Snowflake. - **Composable selectors:** Reference a named YAML selector within `--select` or `--exclude`, making it easier to combine reusable selectors with other selection methods. - **Clearer errors:** More internal Python exceptions are now translated into useful dbt compilation and parsing errors, with cleaner default output and fewer mysterious stack traces. - **Latest version pointer:** Set the new `latest_version_pointer_enabled_by_default` flag to `true` and dbt automatically creates a pointer view for every versioned model in your project, always resolving to the latest version, without any per-model configuration. - and [MANY](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2#quick-hits) others ### What’s next: dbt Core v2.0 [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0) is the next major version of dbt. It replaces the Python-based v1.x runtime with the high-performance, Rust-based foundation developed for the dbt Fusion engine, while keeping the dbt framework open source under the Apache 2.0 license. A major-version transition gives us an opportunity to remove deprecated behavior, enforce a more rigorous language specification, and establish a stronger foundation for the next era of dbt. It also means teams deserve a clear, gradual path to get there. That path starts with v1.12. Try the v2 parser. Resolve outstanding deprecations. See how it behaves with your macros, packages, configurations, and project structure. Tell us what works, and, more importantly, what does not. To learn more, join the dbt Core product, engineering, and developer experience teams for a live virtual release recap and Q&A. We’ll cover the most important changes in v1.12, demonstrate the new parser, discuss how the release fits into the path toward v2.0, and answer your questions live. And to everyone who filed an issue, contributed code, tested a prerelease, joined a discussion: thank you. --- --- title: "Model for the token, not the table" description: "We were burning through Gong's API to feed AI. Modeling the transcripts in the warehouse with dbt cut token costs 20x." url: "https://www.getdbt.com/blog/model-for-the-token-not-the-table" date: "2026-08-17" authors: ["Britton Stamper"] categories: ["Insights"] --- # Model for the token, not the table Gong is one of the highest-value AI context sources we own. When internal usage of Claude started scaling, teams wanted to connect Gong to start pulling from our large volume of call history. The shortest path for those requests flowed through a wrapper around the Gong MCP to the Gong REST API. Unfortunately, that also meant that the team was getting raw, unmodeled context that we couldn’t govern and improve. A single customer-history question fans out through an MCP wrapping Gong’s APIs using `list_calls` to collect IDs, burning 8,300 tokens per call transcript or around 240,000 tokens in context for one account history of ~30 calls. Multiplying this across our entire sales department researching multiple accounts is exactly how you redline the Gong API’s per-second and per-day rate ceilings, which hit our cap very quickly. Gong’s API constraints weren’t the real problem. Our core issue was using a transactional API as a high-throughput context layer for AI. Gong and other sources were never designed for these high-volume, operational use-cases. Databases are. **** ## Compression, context engineering, and cost savings We identified the root cause as the same historical problems that data teams face. Instead of going through our traditional data modeling and warehousing layer, users were hitting source data directly, going to Gong every time with the same "_give me all my transcripts, now summarize them_" request. Analyzing MCP connector usage showed us that Claude was reaching directly to sources that we already had in our data warehouse. Fivetran was already funneling Gong into Snowflake continuously, where we could model data so people would never need the raw source. Serving Gong data from the warehouse through the dbt MCP server removes the API ceiling and lets us model the rich, but often verbose context into exactly the shapes an agent needs. We could even benefit from the broader data warehouse and join Gong data to everything else we know about the account, creating a rich context layer for our team as they used agents like Claude to help them work. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/9788483fe10edcc2adb0a37bf301444b679579c6-2480x1520.png) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/cb1bb2a01332ce4a865edd12be3fe0f400600a0a-2120x1432.png) ## Higher-quality context at lower cost By leveraging our warehouse and dbt, we took care of our API constraint issue _and_ opened up significant potential token savings. Ingesting the raw Gong data into our data warehouse lets us use data modeling to summarize each call, eliminating noise and extracting only relevant signal. Given how valuable Gong data is at informing our business and how frequently it’s used, it made sense to set up pipelines on top of it that we could use for all the call transcript use-cases. Actively compressing transcripts shrank data volume by 20x or more, slashing a 60-minute call from 10,000+ tokens to just a few hundred. The same transcripts, summarized, consume 74x fewer tokens and are now 99% cheaper to serve to AI agents while preserving or improving context quality. We can run hundreds or thousands more Claude sessions a day at the same cost, making ROI skyrocket. The compression itself runs as an incremental dbt job as data enters the warehouse, with inference priced through Snowflake credits, Vertex batch pricing, or Databricks AI Function pricing, all materially cheaper than interactive API calls and paid once per call rather than every turn. The result is order-of-magnitude savings: the cost of compressing 1,000 calls in batch is in the low tens of dollars. Because compression is a one-time cost per call and the savings repeat on every read, the batch job pays for itself within the first meaningful agent session. ## What token efficiency looks like at enterprise scale _Based on Claude Sonnet 5 pricing as of August 2026, subject to change: $3 per million input tokens, $15 per million output tokens._ ## The pattern This pattern is exactly the work dbt is built to do: model the transcripts in the warehouse, compress them with warehouse-native AI functions, and serve the compressed form to agents through the dbt MCP server. We just needed to dogfood things internally to show that dbt transforms data for AI as effectively as it does for BI: 1. **Fivetran delivers the Gong transcript** into the warehouse when a call concludes using an existing connector, no new infrastructure required. 2. **A dbt model invokes a warehouse-native AI function** (Snowflake Cortex COMPLETE, Vertex AI batch prediction, or Databricks AI Functions) to produce a structured summary, capturing deal-relevant points, named entities, sentiment, action items, and objection types. A 10,000-token transcript typically compresses to 500-to-1,000 tokens of structured summary without losing the signal an agent needs for deal context. 3. **The dbt MCP server serves the compressed form** to the asking AI client. This is the same connection AI clients already use for any other dbt-modeled data, and the summary looks like just another well-modeled table. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/e19fcd728d5fb653a3c4da3eb92a7f40cf4aa9f3-1650x578.jpg) Gong was our initial test case, but the pattern generalizes to any token-heavy, durable data source. The same architecture works for engineering context from sources like email archives, Slack channel logs, support tickets, contracts, or marketing collateral. ## Three ways to serve Gong Head-to-head, here’s how using the official Gong MCP, an MCP wrapping Gong’s APIs Gong MCP, and the dbt MCP server stack up against each other for serving unstructured Gong data as AI context: **** ## Cost comparison: Gong vs MCP wrapper vs dbt There's no single best tool for all queries, but there _is_ a best tool per question type. Gong’s official MCP is efficient when all you want is a quick, server-synthesized, single-account summary; an MCP wrapper/gateway platform is good for pulling raw transcript text. Everything else—anything filtered, aggregated, joined, or trended, i.e. the majority of the analytical questions people actually ask—should go through dbt’s modeling and warehousing layer. _Note: Cost calculated using Opus pricing at $5/million tokens, anchored to 8,314 raw and 112 brief tokens per call._ **** ## Get started If you want Gong context wired into your AI workflows, the path forward is to stop connecting directly to Gong's API, partner with the data team to get the Fivetran + dbt + warehouse-AI pipeline modeled. Any AI client (Claude, agents, whatever surface you’re using) connects to the[ dbt MCP server ](https://github.com/dbt-labs/dbt-mcp)as a single governed entry point and gets access to dbt-modeled data through the same surface used for everything else in the warehouse. Analytics engineers have always transformed structured data to build tables for humans and BI tools. Now we can apply the same general principles to unstructured data to build models and engineer context for AI agents. Modeling Gong (and similar sources) for AI is among the highest-impact data-team work available right now, with token efficiency as a first-class design constraint. It’s time to expand beyond thinking what metrics can we serve and start treating call transcripts and other qualitative data as first-class, governed warehouse assets. Now the first question to ask of any source you wire to AI is _what does the raw form cost per question, and what is the smallest modeled form that still answers it?_ --- --- title: "Why agentic projects fail and how to fix them" description: "Why do some AI deployments yield great successes while so many still crash and burn? The answer is in the data." url: "https://www.getdbt.com/blog/why-agentics-projects-fail-and-how-to-fix-them" date: "2026-08-14" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Why agentic projects fail and how to fix them It’s amazing what deltas exist between AI implementations. [Wayfair](https://openai.com/index/wayfair/) built agents to support its suppliers and now automates 41,000 support tickets a month. [C.H. Robinson](https://www.langchain.com/blog/customers-chrobinson) built agents that read shipping-request emails, connect information across messages and attachments, fetch additional context, and automatically create more than 5,500 shipping orders per day, saving 600 person-hours per day. But then there's [Klarna](https://www.customerexperiencedive.com/news/klarna-says-ai-agent-work-853-employees/805987/), which one analyst called the "poster child for bad AI deployments." A customer service agent meant to do the work of over 850 employees contributed to quality issues and declining customer satisfaction. Same technology, wildly different outcomes. And the deciding factor is rarely the model. As we’ll discuss below, what separates the wins from the cautionary tales is the underlying data, and whether users can trust it. **** ## Why agentic AI projects fail without trusted data Agentic AI is undergoing the most aggressive technology adoption curve in a generation. According to the [2026 Gartner CIO and Technology Executive Survey](https://www.gartner.com/en/articles/hype-cycle-for-agentic-ai), more than 60% of organizations plan to deploy AI agents in the next two years, and only 17% have done so today. [McKinsey estimates](https://www.mckinsey.com/capabilities/risk-and-resilience/our-insights/deploying-agentic-ai-with-safety-and-security-a-playbook-for-technology-leaders) that gen AI could add $2.6 trillion to $4.4 trillion in value annually across enterprise use cases. Yet [Gartner also projects](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) that more than 40% of agentic AI projects will be canceled by the end of 2027. A [2026 Fivetran report](https://www.fivetran.com/resources/reports/the-2026-agentic-ai-readiness-index) found that only 15% of organizations are fully ready, even as the vast majority have already invested millions. That gap between ambition and readiness has a cause, and it's a specific one. The key limiting factor for successful agentic AI implementation is generally not model quality, but poor data quality and governance. Frontier AI labs have released powerful foundational models. But if the data that feeds them is inaccurate, incomplete, or inconsistent, you get poor results. [Agentic AI](https://mitsloan.mit.edu/ideas-made-to-matter/agentic-ai-explained) extends the reasoning ability of generative AI to decisions and actions performed through software, not only producing information but also acting in the world. That shift is exactly what raises the stakes. Without adequate data and context, agentic AI can misanalyze a situation, choose the wrong response, perform the wrong action at scale, and cause cascading workflow errors. With poor security, governance, and accountability, including at the level of data assets, it can become difficult or impossible to trace the origin of a wrong decision and remediate it. The public examples are instructive. A chatbot deployed by [Air Canada](https://www.theguardian.com/world/2024/feb/16/air-canada-chatbot-lawsuit) gave a customer inaccurate information about bereavement pricing, failing to refer to and correctly cite the company's internal policies. [Replit's software development copilot](https://fortune.com/2025/07/23/ai-coding-tool-replit-wiped-database-called-it-a-catastrophic-failure/) deleted a production database during a code freeze and created false data in the process, the kind of incident that hard production controls should prevent, whether the actor is human or AI. We find it useful to organize the risks posed by agentic AI into three categories: 1. **Operational correctness risk:** the agent acts on stale, missing, or misunderstood data. 2. **Control-plane risk:** the agent has the wrong permissions, weak auditability, or poor security boundaries. 3. **Human-system risk:** humans overtrust, under-review, or cannot effectively supervise the agent. The through-line in all three is context. For agents, stale data is not merely an analytics problem; it can become an operational action taken on the wrong version of reality. A dashboard built on last week's numbers is a bad report. An agent acting on last week's numbers is a bad decision, executed automatically, at machine speed. **** ## Agents demand more from your data than dashboards ever did Agentic AI imposes far greater demands on an organization's data infrastructure than human-centric analytics workflows, especially reports and other forms of decision support. A human analyst consumes data intermittently; an agent consumes it continuously. A human can absorb tacit, tribal knowledge through experience and can intuit, remember, or investigate where a data asset came from. An agent needs explicit access to context, explicit governance, and access to data lineage. The key challenge is the lack of a governed, consistent context across the enterprise data estate. AI agents need more than raw data. They need reliable data movement, shared business logic, semantic context, lineage, and access controls across every source and consumer. Meeting that requirement rests on two pillars: automation and centralization. Centralizing data and ensuring that it's inventoried and defined once is essential for scaling access to trusted data, controlling infrastructure costs, and managing compliance risks. Once data is centralized, your team has to systematically transform, that is, model, it into a usable context layer for AI. That's what makes lineage, semantic layers, and governance critical: - [**Lineage**](https://www.getdbt.com/lp/data-lineage) shows users, including agents, where data came from and whether it can be trusted. - [**Semantic layers**](https://www.getdbt.com/product/semantic-layer) apply shared business meaning to tables, granting humans and agents alike a shared, consistent understanding of how data maps to real-world business concepts. - [**Governance**](https://www.getdbt.com/lp/data-governance) controls what users and agents can access, decide, and change. There's one more requirement that's easy to underrate: interoperability. As AI tooling continues to evolve, the best model, compute engine, orchestration layer, or activation channel for one workflow may not be the best for another. Interoperability gives teams the freedom to connect systems without duplicating data, rebuilding pipelines, or locking agent workflows into a single vendor. If each AI use case requires copying data into a proprietary silo, organizations lose governance, portability, and control. ## Where agentic AI actually works best Deciding where to point an agent matters as much as the infrastructure underneath it. AI has a very jagged ability profile due to its design, excelling at some tasks while deficient at others. It's strong at pattern recognition and completion for text and code, including drafting, editing, summarizing, translation, and coding. It's strong at ideation and brainstorming, especially when breadth is required, and the cost of a bad suggestion is low. It’s skilled at reasoning through problems with clear, specified premises and constraints. On the other hand, it's weak at discerning truth from plausibility when facts are obscure or highly specific and at knowing when not to answer. It struggles with open-ended problems that require causal reasoning from limited evidence and long-horizon planning, and with adversarial interactions. That profile points to a clear set of characteristics for good agentic use cases: - High volume - Repeatable structure - Text/code-heavy inputs - Clear success criteria - Low-cost human review - Reversible or low-risk actions - Available authoritative data Poor use cases have the inverse: ambiguous accountability, high legal or safety stakes, sparse data, adversarial users, long-horizon planning, and irreversible actions. It helps to remember that these systems tend to augment work rather than replace it. In 2016, [Geoffrey Hinton](https://www.nytimes.com/2025/05/14/technology/ai-jobs-radiologists-mayo-clinic.html), who would later win the 2024 Nobel Prize in Physics for his work on artificial neural networks, predicted that radiologists would be extinct as a profession by 2021 due to AI image recognition. By 2025, radiologists' pay, employment, and workloads had never been higher. AI is far likelier to augment complex workflows than eliminate roles. We've put these principles to work ourselves. Fivetran's Chief Product Officer uses agentic AI to perform conversational analytics on Jira data. Directly querying Jira's MCP server was untenable at scale, so the team moved Jira data into BigQuery via Fivetran, used a Claude Skill to query it, and produced product-ops insights in hours rather than multiple analyst sprints. Separately, our support team embedded a custom AI app in Zendesk to answer questions, draft responses, summarize handovers, and find similar tickets. Built using Fivetran and dbt, it centralizes knowledge from Zendesk, Slab, Jira, GitHub, Google Drive, Gong, Salesforce, and docs. The common thread: each agent is narrow, high-volume, and grounded in authoritative data it can‌ reach. **** ## Build an agent you can trust Once you've picked a workflow, the build itself is more approachable than most teams assume. Building agentic AI models from scratch is a complex undertaking that can cost many millions of dollars and months of development time. A more practical and less risky option is to augment a foundation model with your organization's unique, proprietary data using a RAG architecture. You can create specialized agents that perform specific tasks by interacting with your operations through the [dbt MCP server](https://www.getdbt.com/blog/mcp) and similar controlled interfaces. The safest way to roll that out is in tiers of progressively growing autonomy: - Start with read-only agents that retrieve and summarize information. - Then, build drafting agents that prepare outputs for human review and final implementation. - Next, build bounded write-back agents that act within strict, narrow limits. Potentially risky actions should require approval, while sensitive, irreversible, regulated, or safety-critical actions should remain prohibited from autonomous execution. Getting each of those tiers right depends on several details, such as your reference architecture, the discipline of [context engineering](https://www.getdbt.com/blog/bringing-structured-context-to-ai-with-dbt), and a concrete readiness checklist. We’ve laid out this blueprint in [The data leader's primer for agentic AI](https://www.getdbt.com/resources/the-data-leaders-primer-for-agentic-ai). It walks through the reference architecture step by step, the strengths-and-weaknesses map for choosing use cases, and the checklist we use to take an agent from idea to production so your project avoids ending up part of the 40% cancellation statistic. [**Download the full guide**](https://www.getdbt.com/resources/the-data-leaders-primer-for-agentic-ai) for the roadmap. --- --- title: "How dbt State cuts warehouse compute and speeds up every run" description: "How Fanatics cut warehouse compute by only rebuilding what's changed" url: "https://www.getdbt.com/blog/dbt-state-use-case" date: "2026-08-14" authors: ["Daniel Poppy"] categories: ["Product"] --- # How dbt State cuts warehouse compute and speeds up every run Every day we see about 859,000 dbt jobs run. Most of them run hourly. And most run hourly even when the data hasn't changed, so thousands of models get rebuilt to produce the same results. For a long time, that made sense. Cloud compute kept getting cheaper per unit, and it was easy to just use more of it. Then everyone started using more data: more dashboards, and thus more pipelines powering machine learning models and recommendation engines. Once agents arrived, anyone could write a complex SQL query, or even create a dbt model. What started as a reasonable way to run your pipelines got expensive. Those cheap cloud costs started to skyrocket. [dbt State](https://www.getdbt.com/product/dbt-state) is our answer. It takes what was a stateless engine, dbt, and makes it stateful: every time you run `dbt run` or `dbt build`, it only builds what's‌ changed and skips everything that hasn't. The tagline we keep coming back to is simple: build what's changed, skip what hasn't. Here’s how it works and how one of our customers uses dbt State to push model reuse rates as high as 25%. **** ## Why we built dbt State For all of our data people, this was always about more than compute. Saving 30% on the data warehouse bill is great, and we wouldn't turn it down. But the thing you feel every day is the time, the focus, and the trust it takes to keep a project running. Think about the friction: - Every time you develop a new model, you have to ask which selectors it belongs in. Managing 20 different selectors across 10,000 models becomes a problem in itself. - Every time you run in development, you wait for 100 or 200 models to build, 10 or 20 minutes, just to figure out whether you built the right thing. That costs focus. - As costs climb and the business doesn't see more value, because the data isn't refreshing any more often, people start asking whether it's worth it. dbt State started at our last [dbt Summit](https://www.getdbt.com/dbt-summit/), where we shipped [state-aware orchestration (SAO)](https://docs.getdbt.com/docs/deploy/state-aware-about) into preview. The feedback was clear. Customers were saving 30%+ on their data warehouse bill. They came back with two questions. First, why is this only in production, when I'd love to use it in staging, [continuous integration (CI)](https://www.getdbt.com/blog/adopting-ci-cd-with-dbt-cloud), or development to make my runs faster? Second, why is it available only in the dbt platform on the [dbt Fusion engine](https://www.getdbt.com/product/fusion) when I want to run it anywhere? We listened. We took the best parts of SAO and made them work everywhere. The result is dbt State. ## What dbt State does The basic idea stays the same as SAO: build what's changed, skip anything that hasn't. That unlocks three things. - **Optimize costs.** Teams see 30% less warehouse compute on average by only building what's changed upstream and adhering to the freshness [service-level agreements (SLAs)](https://docs.getdbt.com/reference/resource-configs/freshness) you set, so the business isn't paying for fresh data more often than it uses it. - **Run freely.** When you kick off a `dbt run` or `dbt build`, you can be confident you won't build 100 models when only 10 have changed. You'll build 10. - **Speed up development.** You can increase the speed of every run without resorting to custom workarounds like sampling or using different data across development and production. You turn dbt State on, and it just works. ## Run it anywhere you run dbt dbt State is available natively in the dbt platform and locally in [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion). If you're on v2.0, you can run `dbt login` and set up dbt State right away. If you're on v1, from v1.7 through v1.12, it's available as a plugin: run `pip install dbt-state` and you can be up and running across development and production in less than five minutes. There's no platform upgrade required. Anywhere really does mean **anywhere**. Run dbt in an orchestrator like [Dagster](https://dagster.io/) or [Airflow](https://airflow.apache.org/). Add it to your own scheduled setup in a Python environment on EC2. If you can run dbt, you can run dbt State. That last point is the part that matters most. SAO and dbt State were built on the same premise, skipping builds when the upstream data or code hadn't changed. What we added with dbt State was reach: no longer requiring Fusion, no longer limited to production, and a way to reuse data you've already built instead of paying to rebuild it. When the same logic and data already exist somewhere else, say in a staging or production environment, dbt State clones them in instead of rebuilding, and you won't even notice the difference. ## How dbt State works It helps to separate two things: the **data plane**, where you run dbt, and the **control plane**, where dbt State runs. In the data plane, every `dbt run` executes against your warehouse, where it has access to your data and builds your models. The control plane is where dbt State decides what actually needs building. It keeps a list of your tables and a hash of the last data state and code state. There's no actual data there, just a hash of what the state was. So on every run, dbt State asks a short series of questions for each model. Has it actually changed? If it has, does a version already exist somewhere else that we can clone in? Only when the answer to both is no do we build it. [**For a demonstration of dbt State in action, check out this on-demand webinar.**](https://www.getdbt.com/resources/webinars/dbt-state-build-what-s-changed-skip-what-hasn-t) ## Paying only for what you reuse All of this starts with a 30-day trial for every user. Run `pip install`, try it in development, run it in production, and tell us what you find. When you do start paying, we only want to charge you for the parts that produce value. Our metric is the daily active target table (DATT): a model or test that dbt State reused. We count unique reuses per day, so we don't double-charge you. If your `fact_orders` and `dim_customers` models rebuild six times a day but the data didn't change, you're not paying for six reuses each. You're paying for two, one per unique table, no matter how many times it was reused that day. ## How Fanatics Betting and Gaming put dbt State to work One of the teams that turned dbt State on early was [**Fanatics Betting and Gaming**](https://www.fanaticsinc.com/fanatics-betting-gaming), and they saw double-digit compute savings on the very first project they tried it on. Alvin Chai, a senior analytics engineer at Fanatics, runs a team responsible for building and maintaining gold-layer data models across the betting and gaming business, everything from sportsbook trading analytics and product insights to regulatory and financial reporting. Their stack was [Snowflake](https://snowflake.com) for the warehouse, dbt Core, and Airflow to orchestrate. About 18 months ago, they moved to the dbt platform to remove a bottleneck: refreshes and backfills used to require a specific data engineer to kick them off. Today they run about 11 projects, one per analytics domain, 8 of them on Fusion, with a couple thousand models in total. Democratizing scheduling gave every analytics team the freedom to schedule their own refreshes. The downside, as Alvin says, was overscheduling with little visibility into it: "Maybe you have an analytics team that might set up their models to run every 30 minutes. But the source that model is built on might run maybe only every hour or so." dbt State surfaced their model reuse rates and effectively put guardrails on overscheduling, refreshing models only when they actually need it instead of rebuilding the same output over and over. That made the whole thing more hands-off, so the team no longer had to pester other analytics domains about their costs getting too high. The rollout is a useful lesson in how to get value out of dbt State. Fanatics started with lower-risk work. Their pilot was a governance project supporting responsible gaming and anti-money laundering use cases, chosen because it had few downstream dependencies through [dbt Mesh](https://www.getdbt.com/blog/data-mesh-architecture-explained) and would be easy to roll back. They then worked down the hierarchy by risk, migrating everything except their most critical regulatory and financial reporting. Turning dbt State on with no configuration gave them a model reuse rate of about 0.2%, essentially nothing. The effectiveness came after they configured source freshness and declared model-level SLAs, what we now call lag tolerance. Their approach was to set a baseline model-level SLA for every model based on its current refresh frequency, add a more aggressive project-level SLA on top, and write a small script to translate their existing schedule tags into those SLAs. A model tagged "daily," for example, became a 24-hour refresh window. That took the pilot from 0.2% to roughly 15% model reuse. Historically, overscheduled projects later hit as high as 25%. Others landed nearer 5%. Across all their jobs, the average sits around 8%. The compute savings are real, but Alvin's take is that the biggest long-term win is operational simplicity: "The question kind of turns from, 'What job should this model belong to?' to 'How fresh does this model actually _need _to be?'" Instead of juggling separate jobs by cadence and managing tags for each one, Fanatics can move toward DAG-based jobs: point at a final consumption model, refresh everything upstream of it, and let the model-level SLAs handle the different freshness requirements underneath. The payoff is where the team spends its time. Rather than burning mental energy figuring out how to optimize schedules, they can focus on building new models and delivering more insight and business value. ## Try it yourself A decade of dbt has been about bringing engineering discipline to analytics. dbt State extends that discipline to the compute itself: stop paying to rebuild what hasn't changed, wherever you run dbt. You don't have to take our word for it. [**Create a free dbt account**](https://www.getdbt.com/lp/dbt-free-account) or `run pip install dbt-state` on your existing project, and watch your development runs get faster and your warehouse bill get smaller. [**To see dbt State in action, watch the on-demand webinar.**](https://www.getdbt.com/resources/webinars/dbt-state-build-what-s-changed-skip-what-hasn-t) --- --- title: "dbt Summit 2026: the keynotes and product sessions" description: "The dbt Summit 2026 keynotes and the product sessions behind them." url: "https://www.getdbt.com/blog/dbt-summit-2026-keynotes-product-sessions" date: "2026-08-10" authors: ["Daniel Poppy"] categories: ["Learn"] --- # dbt Summit 2026: the keynotes and product sessions There’s something we keep hearing from data teams. A chief data officer at a health insurance administrator framed it well recently. He has two data engineers he wants to fast-track onto AI work. He has executive support. He has budget. And he still can’t start. “Everybody's making noise around AI. Can you do something around AI? And I'm thinking, ‘But your data is still not at the level where you can put an AI agent on top of that.’” He already knows what most organizations are about to find out the hard way. It’s a gap this year’s dbt Summit content happens to address. [dbt Summit lands September 15-18 at The Cosmopolitan in Las Vegas.](https://www.getdbt.com/dbt-summit) ## You have done this before Ten years ago, dbt changed what it meant to be a data professional. The analytics engineer emerged, with production-grade pipelines at speed and scale, governed data, and real influence over decisions across the business. It’s easy to forget how that felt at the time. The shift was disorienting. Plenty of people weren’t sure where they’d land. It turned out to be one of the best things to happen to careers in this field. We’re at that moment again. AI is rewriting how data gets used, and it’s raising the stakes on every data decision made in the next 18 months. The consumers of data are shifting from analysts in a BI tool to agents working continuously, at machine speed, without a human checking every answer. The teams who thrive will be the ones who level up, which is why you need to be at dbt Summit this year. ## Two keynotes, four pillars The [dbt Summit keynote](https://www.getdbt.com/dbt-summit/sessions/keynote-level-up) is built on four pillars, and they line up closely with what we hear directly from data teams and what they are up against. **Level up the engine** with faster parsing, better scalability, and real cost control, wherever you run dbt. **Level up for AI and what your agents consume**. Structured, governed context is the missing piece between your warehouse and an agent whose answers you’d stake a decision on. **Level up with AI and what you consume** with [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) and the agentic workflows around it, so you build and ship faster while keeping control of quality and spend. **Level up the stack** with an open, flexible infrastructure. Your stack stays yours. Thursday brings the [Community Keynote](https://www.getdbt.com/dbt-summit/sessions/keynote-community-keynote) with Grace Goheen and Jeremy Cohen. This is a celebration of the dbt community, how dbt Core and the engine unite under one framework, and how the work continues in the open. ## Go deeper in the product sessions The keynotes are the headline. The product breakouts and roundtables are where you find out how it all works, straight from the product managers who built it. [**Optimizing your runs for lower compute, fresher data, and faster iteration with dbt State**](https://www.getdbt.com/dbt-summit/sessions/optimizing-your-runs-for-lower-compute-fresher-data-and-faster-iteration-with-dbt-state). Most dbt projects rebuild the same models every run, whether anything changed or not. [dbt State](https://www.getdbt.com/product/dbt-state) checks warehouse metadata and model SQL, works out whether the result would actually change, and then builds, skips, clones, or auto-defers to production. Average compute savings run 30%. Reuben McCreanor walks through freshness SLAs, auto-cloning from prod, and the orchestration logic you get to retire, whether you run dbt Core, the dbt platform, or something in between. [**dbt Wizard: your AI teammate for data development**](https://www.getdbt.com/dbt-summit/sessions/dbt-wizard-your-ai-teammate-for-data-development). dbt Wizard is the coding agent built for data, not software. Ani Venkateshwaran and Brandon Thomson walk through its three modes—develop, analyze, and discover—the agent skills framework that makes it extensible, and a live demo of the fully autonomous analytics engineering workflows dbt’s own team has built on top of it. [**Rebuilding dbt Core in the open: faster runtime, adapters, and docs v2**](https://www.getdbt.com/dbt-summit/sessions/rebuilding-dbt-core-in-the-open-faster-runtime-adapters-and-docs-v2). Years of dbt Core v1.x left teams with a Python runtime that slowed on big projects, fragmented adapters, and a docs experience that couldn’t keep up. Hope Watson unpacks the rebuild: a Rust-based Apache 2 engine, a cleaner adapter model on ADBC and the Arrow ecosystem, and parquet-backed artifacts. You’ll see side-by-side parse speed demos and get a real migration path from v1.x through the v1.12 parser on-ramp. [**The semantic layer is dead. Long live the semantic layer!**](https://www.getdbt.com/dbt-summit/sessions/the-semantic-layer-is-dead-long-live-the-semantic-layer) Semantic layers used to be a solved problem you built years ago for dashboards. Then agents showed up. Zach Mandell makes the case that your semantic layer is now the most load-bearing piece of your AI strategy, shows where MetricFlow and the dbt Semantic Layer are heading, and covers how the dbt MCP server turns governed metrics into context your agents can use. This is the session that answers the hallucination problem directly. [**Product roundtable: open data infrastructure in practice**](https://www.getdbt.com/dbt-summit/sessions/product-roundtable-open-data-infrastructure-in-practice-unlock-your-dbt-projects-with-apache-iceberg-and-mesh). Jack Lowery and Anna Lee host a working conversation on unlocking dbt projects with Apache Iceberg and mesh patterns. Roundtables are small and discussion-first, so come with the problem you’re actually stuck on. [**Apache Ossie: Realizing Semantic Layer Portability**](https://www.getdbt.com/dbt-summit/sessions/apache-ossie-realizing-semantic-layer-portability). Define a metric in one tool, redefine it in the next, and your dashboards, notebooks, and agents all end up disagreeing on what “active customer” means. Apache Ossie, the vendor-neutral, open-source semantic interchange spec co-led by dbt Labs, Snowflake, Salesforce, BlackRock, and RelationalAI, changes that contract. You’ll see how MetricFlow, now open-sourced under Apache 2.0, compiles governed metrics into a portable format so BI, conversational analytics, and AI agents finally agree on the same business logic. [**The next wave of data infrastructure is in the Lake**](https://www.getdbt.com/dbt-summit/sessions/the-next-wave-of-data-infrastructure-is-in-the-lake). The modern data stack got you here, but it won’t get you to where agents need to operate. Russell Christopher and Casey Karst lay out Open Data Infrastructure: land data in open formats like Iceberg in storage you control, then pick the right compute for any job and swap tools without rebuilding pipelines. You’ll leave with a framework for building open, portable, AI-ready infrastructure with Fivetran + dbt. [**We said I do. Now, we’re joined at the DAG**](https://www.getdbt.com/dbt-summit/sessions/we-said-i-do-now-were-joined-at-the-dag). Where did this data come from, what happened to it, and who’s using it? Roxi Pourzand shows how Fivetran and dbt unify ingestion, transformation, and consumption into a single lineage graph—one that falls out of the systems doing the work instead of being stitched together after the fact—unlocking real-cost visibility tied to usage, PII that stays governed downstream, and agents that reason with the full picture. [**With great context comes great autonomy: leveling up your agent context**](https://www.getdbt.com/dbt-summit/sessions/with-great-context-comes-great-autonomy-leveling-up-your-agent-context). An agent querying your warehouse shouldn’t be guessing what “active” means or which revenue table to trust. Ben Moser and Kevin Kim demo how dbt structured context and the Fivetran context layer turn raw metadata, governed semantic models, and real usage into the agent schema that powers accurate results—and how to let agents reason beyond fixed definitions when the question calls for it. More sessions are landing between now and September. Browse [the full agenda](https://www.getdbt.com/dbt-summit/sessions) to build the rest of your week. ## Come find out what you get to build next The organizations that win the AI era won’t be the ones with the best models. They’ll be the ones whose data foundation was ready when it mattered. That foundation already has a name, and you built it. dbt Summit 2026 September 15-18, 2026 The Cosmopolitan, Las Vegas Your next level starts here. [Register for dbt Summit](https://www.getdbt.com/dbt-summit/registration). --- --- title: "From analytics engineer to context engineer" description: "First in a series on the shift from modeling data for dashboards to modeling context for agents. We start with our own Gong data." url: "https://www.getdbt.com/blog/from-analytics-engineer-to-context-engineer" date: "2026-08-06" authors: ["Britton Stamper"] categories: ["Insights"] --- # From analytics engineer to context engineer Most companies are taking shortcuts in their enterprise AI rollouts today, and falling into a classic trap their data teams are already deeply familiar with. Taking the “easy route”—wiring AI straight to MCPs provided by vendors, connecting them directly to source data providers—is collectively costing companies billions in token consumption, vendor API charges, and lack of data ownership and portability. At dbt Labs, we initially were no different. We used vendor MCPs until we discovered how significant these hidden costs actually were. That's when we realized that the most efficient, scalable way to connect AI to data was what we’ve evangelized all along: to centralize our data and engineer context directly within our data warehouse, giving us a massive opportunity to own our AI’s context layer and expand our data team’s remit. ## How it started As Claude rolled out in our organization, one of the top user groups was sales connecting to Gong, Salesforce and other rich context sources to analyze their deals. When our GTM teams requested customer insights to prevent churn and identify high-value accounts, of course we agreed and allowed them to analyze Gong data by connecting AI agents directly via the native MCP. With hundreds of salespeople using this data many times a day, costs added up quickly. Analyzing a single call consumed up to 50,000 tokens; that's $0.25 in token cost just for AI to read a single transcript. We had 500M+ tokens worth of Gong transcript data that we could tap into. With many salespeople running many Claude sessions throughout the day, our total AI spend went up significantly at scale. We knew we needed a new approach. That’s when we decided to ingest the raw Gong data into our data warehouse and used data modeling to summarize each call. By actively context engineering the transcript data to eliminate noise and extract signal, we shrank data volume by 20x and slashed token consumption on a 60-minute call from tens of thousands to just a few hundred token. The question was, if we can transform any of our data into more meaningful context, how could we engineer the smallest, most meaningful context layer that still answers most queries? ## How dbt reduced token costs When our data team had previously modeled Gong data for BI, the ROI wasn't there. Now, though, Claude and ChatGPT agents unlocked new use cases that let us do qualitative data analysis at scale, and the call transcript table became one of the highest value assets in the data warehouse. Sourcing context through the data warehouse is vastly more efficient than pulling it straight from MCPs. Serving that call transcripts from warehouse summaries dropped token costs by roughly 98% while maintaining or improving context quality. In this blog series, we will share core patterns for efficiently modeling trusted context for AI, an end-to-end technical walkthrough of our Gong pilot, the fundamentals of reading data once to serve every agent, and why you can transform traditional analytics engineering into context engineering without changing your stack. ## Qualitative data analysis at scale with engineered context Before AI, working with text data at meaningful scale required either simple and ineffective regex techniques like keyword matching or advanced statistics and data science. Maintaining complex transformations for qualitative data sources like call recordings or support tickets was just not feasible for data teams. As our Gong breakthrough shows, though, we have crossed a threshold where previously ignored data and data sources can now form some of the most useful and valuable agent context. Teams can build models and highly effective context layers for AI, applying the general principles we pioneered with analytics engineering as long as the techniques are applied with AI in mind. Analytics teams have always modeled quantitative data for BI. Context engineering is just modeling data for agents, and that data is far more than metrics: ## A dashboard needs metrics. Agents need context. **Untapped qualitative data is where AI truly pays off.** Instead of just serving structured data to dashboards, data teams can now use unstructured data like PDFs and JIRA tickets to engineer context. LLM advances make it possible to move, store, model, and process these files within the data warehouse and use them as agentic AI context. This is a new way of thinking for a lot of analysts because, historically, data organizations have been allergic to unstructured data like transcripts, and for good reason: our technology was not built to support it. We were forced to change because we were running out of Gong API calls. Everyone was asking very similar questions on very similar data, going right to Gong saying _give me all of my transcripts, now summarize each transcript_. The direct MCPs circumvented our entire traditional data modeling and warehousing world to repeatedly query raw data, which created redundant token costs and caused API constraints. This was when we realized _hey, you know, we actually already have this process whereby we model data into trusted, governed, more useful forms. Why don’t we apply that for data for AI?_ Following our merger with Fivetran, we realized [that moving and modeling unstructured data is already in our wheelhouse.](https://www.fivetran.com/blog/ai-requires-unstructured-data-to-unlock-its-full-potential) Platforms like Salesforce, Zendesk, or Gong now provide critical business context. Modeling this qualitative data and caching pre-built AI summaries in the warehouse delivers reliable, deterministic context while eliminating multiple token-heavy MCP calls, tool proliferation, and API constraints. Users seamlessly access this modeled data through a single context connector in the context layer via MCP. This drastically lowers token costs, simplifies the user experience to a simple chat, and opens up reusable, community-driven data patterns across the entire organization. ## How to do cost-effective context engineering There’s no magic in this approach. Data teams turn raw data, whether quantitative, qualitative, or semi-structured, into agent context through the same processes they already practice without materially changing the stack: - **Compress:** reduce a large corpus to the smallest forms that contains only what’s relevant - **Enrich:** join data to other relevant information so that it’s easier to access everything needed, like opportunity details and qualitative deal history. - **Describe:** provide information around what the data means and when it’s applicable - **Govern:** decide exactly what each agent is allowed to see to do its work ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/66f6a0a0ab2a94cd80765e8c857b7cf60ea679fb-1564x434.jpg) The data team models the context once, and every tool and agent reads from the same trusted context layer. The analytics engineer is now the context engineer. It’s the same skillset, except beyond modeling metrics you’re also mapping the entire data estate and every data source is now in play. All of this happens without materially changing your data stack. The data team builds a [structured shared context layer](https://www.getdbt.com/blog/bring-structured-context-to-conversational-analytics-with-dbt) for the entire organization, functionally layering AI over your original data stack in the same way that dbt is currently layered over your data warehouse. The whole company is now the consumer because every team and the agents they use can access this context layer, not just people who write SQL queries. Functionally, though, how do you move, transform, and manage structured, semistructured, and unstructured data into the context layer? 1. **Getting data, both traditional tables and unstructured files, into the data warehouse **(Snowflake, BigQuery, Databricks) is Fivetran's job. If tabular data needs further preparation (filtering, denormalizing, joining, and aggregating), then dbt allows the creation and execution of that transform logic in an open and portable way. 2. **Automating pipelines that utilize the SQL-based AI functions of your destination** is dbt’s job. For example, Snowflake Cortex provides SQL functions such as AI_EMBED, AI_COMPLETE and AI_PARSE_DOCUMENT that can be executed as part of, and orchestrated by, dbt models. BigQuery and Databricks offer similar functionality. Raw text fields, structured data from files like spreadsheets, and even replicated unstructured files like PDFs become queryable, retrievable context available to both human data users and AI agents in the shared context layer. The documents and data sources they seek information from are already in the same platform, moved and indexed automatically by Fivetran and processed as part of your dbt models. ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/1543641da2fcc03b3228ed3497d1e432f7c8b4b8-1322x618.jpg) The stack beneath the context layer remains the same: Fivetran moves your data; dbt models it. Now, though, this includes the unstructured context that makes AI systems powerful. All the Salesforce email bodies, Jira ticket attachments, call transcripts, and contract PDFs that have always existed but never made it into a pipeline are now fully referenceable, trusted context. ## Best practices are beginning to emerge We are continuing to build the vocabulary, packages and open source projects that analytics engineers will need to use to do context engineering. We’ve come up with a few best practices we can share, with much more to come. Here’s a preview of some our team uses: - **Don’t aggregate context, generate it: **Use warehouse-native AI to create new context by parsing, chunking, embedding, transforming and joining related information into modeled data objects that agents can query. Aggregating can cause context to get lost, so maintaining the context quality through the whole pipeline is critical. - **Context layer architecture:** One shared, structured and governed, open-format canonical context layer that every engine can write to and every agent can read from. The same single source of truth the data industry has always focused on is now pointed at AI as the primary consumer. - **Read once, write many: **Reads are the most expensive part of context engineering. Do the expensive read with an LLM once, in batch processing where it’s cheaper; then serve the modeled form to every agent in the context layer - **Incremental context maintenance: **Agents do actions, and they need correct, current information to act on. Context can’t simply be snapshotted, it must be updated as new events and information come in. Incremental models that update context’s current state (with pipelines built to capture the state changes so that they are auditable) are critical. Until recently, the context path of least resistance was to just take the raw API and connect the MCP server. Now, data analysts can step in and say, "Use our existing data. We'll augment that with some of the AI capabilities you're asking for, process it once, and make it accessible to everybody to use an unlimited number of times.” Now, the data team owns AI and becomes the mission-critical team for the next era of businesses. [Fivetran + dbt Labs are building the data foundation for agents you trust. Join us at dbt Summit, where data practitioners and leaders come together to shape the future of data and AI.](https://www.getdbt.com/dbt-summit) --- --- title: "Retiring the dbt Snowflake Native App" description: "The dbt Snowflake Native App retires in November 2026. Here's what it means for you." url: "https://www.getdbt.com/blog/retiring-the-dbt-snowflake-native-app" date: "2026-07-24" authors: ["Kyle Dempsey"] categories: ["Product"] --- # Retiring the dbt Snowflake Native App We’ve made the decision to retire the **dbt Snowflake Native App** from the Snowflake Marketplace. The app will enter maintenance mode in **July 2026** and will be fully removed in **November 2026**. If you are one of the small number of customers currently using the Native App, your dbt Labs account team will reach out directly to support a smooth transition. For the vast majority of dbt + Snowflake users, this change has no impact on your workflows. ## What this means for you ### If you don't use the native app No action is required. This change does not affect dbt platform, dbt Core, the dbt Semantic Layer, or any other dbt product or integration. ### If you currently use the native app For the small number of customers with an active installation: - **Your app will continue to function through November 2026.** - **During the maintenance window** (from July 2026 to November 2026), we will provide critical bug fixes and security patches only. No new features will be released. - **Your dbt Labs account team will contact you directly** to walk through your specific situation and discuss your options. If you do not hear from your account team, please email [support@dbtlabs.com](mailto:support@dbtlabs.com) by **July 2026**. - **After November 2026**, the app will be fully delisted from the Snowflake Marketplace, and access will be discontinued. ## Why we're making this change When we originally launched the dbt Snowflake Native App, our goal was to bring dbt-powered AI capabilities directly into the Snowflake environment. Since then, the ways dbt and Snowflake work together have evolved significantly — and customers now have access to better-supported, more capable options. - **dbt’s Semantic Layer is available in Snowflake today — without the Native App.** Customers can model governed metrics in dbt and use them in Snowflake through Snowflake Native integrations (e.g., Semantic Views / Snowflake Intelligence) with dbt providing the trusted semantic definitions behind the experience. If you were using the Native App to bring dbt context into Snowflake, these options deliver a deeper, more reliable integration. - **The AI landscape has matured.** When the Native App launched, bringing a dbt-powered chatbot into Snowflake was a reasonable experiment. Today, Snowflake Intelligence and Cortex provide a far richer AI experience — with dbt's Semantic Layer powering the trusted context behind them. - **We want to focus investment where it benefits you most.** Rather than maintaining an app that hasn't kept pace, we're directing our partnership efforts toward the integrations that customers are actually adopting and finding valuable — like Semantic Views, OSI, and native dbt execution on Snowflake. dbt will continue to support Snowflake data teams with interoperable products and strong integration. We’ll also continue to help teams deliver transformations while providing observability into AI agentic data operations — so customers can build and operate trusted, governed data products that power analytics and AI on Snowflake. ## Support and feedback If you have questions about this change or need help transitioning: - **Email:** [support@dbtlabs.com](mailto:support@dbtlabs.com) - **Your account team:** Your CSM or account executive can help with your specific situation We welcome feedback on the timeline. While the decision to retire the Native App is final, we're open to adjusting dates if you need additional time. Please reach out if the current timeline creates a hardship for your team. --- --- title: "Fivetran + dbt Labs: The future of dbt Core v2.0" description: "AI has changed the game. dbt is changing along with it. Learn more about dbt Core v2.0 and the future of dbt itself." url: "https://www.getdbt.com/blog/fivetran-dbt-20-future" date: "2026-07-21" authors: ["Daniel Poppy"] categories: ["Product"] --- # Fivetran + dbt Labs: The future of dbt Core v2.0 Last year, we made two announcements that injected a bunch of uncertainty into the dbt community: [the dbt Fusion engine](https://www.getdbt.com/product/fusion) and [coming together with with Fivetran](https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means). Change is scary, and the questions came fast. What was the future of dbt Core? One Reddit user went so far as to predict the "final nail in the coffin of OSS dbt," shorthand for open-source software. What has‌ turned out to be true over the past year is that we have wanted to ship more code in the open, not less. And we anticipate that being true for many years to come. The clearest proof is [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion): the same fast, capable foundation that powers the dbt platform, now under an Apache 2.0 license. Alongside it, we've shipped [dbt State](https://docs.getdbt.com/docs/deploy/dbt-state-about), a caching layer that cuts customers' dbt-driven compute by 30%+, and dbt Wizard, a coding agent purpose-built for dbt. ## Why the data stack needs to be rebuilt for agents Taylor Brown, Co-founder and COO, Fivetran + dbt Labs, and Tristan Handy, Co-founder and President, Fivetran + dbt Labs, have been working for a long time to bring the two companies together. The two joined a webinar recently to discuss the rationale, and share what comes next. For the full conversation, including the live Q&A, [watch the webinar recording](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). And if you want to see the speed for yourself, [create a free dbt account](https://www.getdbt.com/lp/dbt-free-account) and invoke dbt Core v2.0. Back in the 2010 to 2014 timeframe, everyone was trying to solve the problem of getting data out of siloed places and leveling up the overall analytics infrastructure to drive business value, largely for reporting. That's where Fivetran plus dbt originally helped coin and build the modern data stack. Over 100,000 teams adopted dbt as a standard for transformation. Over 8,000 customers adopted Fivetran for data replication. [The age of AI has changed the game](https://www.getdbt.com/blog/how-ai-will-disrupt-data-engineering). As Taylor puts it, the outcome that data stacks drive is no longer just for humans and reporting. AI and agents are now large consumers of data, and they're driving business operations and revenue applications. That shift exposes three structural breaks in the stack at agent scale: - **Trust breaks.** If you don't have complete data, consistent metrics, and lineage, you don't have trust. And trust is now infrastructural, because agents are taking action on this data, not just humans who understand its quirks. - **Scale stalls.** With ungoverned data, fragmented context, and locked-in formats, agents will amplify how bad the underlying data is and operationalize it in negative ways. - **Cost explodes.** Repeated retrievals and unoptimized modeling leave you with many more agents running on very expensive compute. When we hit the cloud era, the cloud was so much more useful than on-prem systems that we had to rip everything up. Fivetran and dbt are a result of what happened there. We're doing that type of work again in the agentic era because we're going to see at least an order-of-magnitude increase in data consumption. Our answer is an architecture we call [Open Data Infrastructure](https://www.getdbt.com/blog/what-is-open-data-infrastructure): a data stack built for agents and humans that's flexible by design, fresh and trusted by default, and efficient at scale. An open data infrastructure is: - Built on open standards, separates storage and compute. - Loads into modern data lakes in open formats like Iceberg and Delta. - Relies on strong data movement and transformation in the pipeline layer. - Provides rich metadata and context in the management and governance layer. Companies are already putting this together. - [**Zendesk**](https://www.zendesk.com/), the global customer experience platform, [used Fivetran plus dbt to scale out analytics agents and AI across the entire enterprise](https://www.getdbt.com/resources/coalesce-on-demand/coalesce-2025-how-zendesk-built-a-cross-domain-multi-platform-data-strategy) in a fraction of the time it would normally take. - [**Shutterstock**](https://shutterstock.com) built a trusted, more real-time analytics and AI platform for emerging AI workflows. - At [**Inova Health**](https://www.inova.org/), a leading nonprofit healthcare provider, [Jon McManus and his team compressed a four-year data modernization roadmap into six months](https://www.fivetran.com/case-studies/inova-health-compresses-4-year-roadmap-into-6-months-to-power-ai), with Fivetran and dbt as key parts of the new architecture. For healthcare, that pace is unheard of. ## dbt Core v2.0: One engine for all of dbt If you've been on this journey with us for any length of time, you've seen us trying to do two things at once. The first is to support a widely used piece of open source software infrastructure. The second is to make money so we can keep doing the maintenance. About 18 months ago, [dbt Labs acquired SDF Labs.](https://www.getdbt.com/blog/dbt-labs-acquires-sdf-labs) Why? Simple: their engine was, in many ways, technically superior to long-term dbt Core. It's written in Rust and includes a bunch of other capabilities. We combined dbt and the SDF engine and shipped it as the dbt Fusion engine. But Fusion and dbt Core lived side by side, on different technical foundations (Rust and Python) and different licenses (Elastic and Apache 2.0). Over the past year, as we did the work of making Fusion ready for general availability (GA), we realized we wanted to bring these two engines together. That work culminated in the launch of [dbt Core v2.0](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion). The alpha shipped on June 1, and it's still early, but the Apache-licensed dbt project is now based on the Rust implementation that Fusion created. The first thing you'll notice is that it's just really fast. The first time you invoke dbt Core v2.0, the responsiveness in your command-line interface (CLI) is dramatically improved. And there's a change we’re especially excited about in how we deal with metadata. dbt has published artifacts like manifest.json and run_results.json for a long time, but dbt now produces Parquet files as well, which form the foundation of a context layer. You can query them locally with DuckDB and immediately find out facts about your dbt project. Fusion isn't going anywhere. Fusion is a superset of dbt Core v2.0. Both are based on the same core technology, and pure OSS dbt Core is now a strict upgrade over v1. Installing the Fusion binary gets you everything in Core plus more, and the heart of that "more" is SQL comprehension. Fusion natively understands the SQL you write. That unlocks developer experience benefits such as easier refactoring, autocorrect, and column-level lineage in your docs. [You can see the full comparison of Fusion and dbt Core v2.0 in our docs](https://docs.getdbt.com/docs/fusion/fusion-availability). The same binary you can use free, without ever speaking to us, now also supports dbt login, which gives you access to proprietary features, some free and some paid. And if you're a dbt platform customer, upgrading to v2.0 or Fusion is straightforward: an auto-migrator tool with AI to fix any remaining issues. All of this lands just past a milestone that's honestly shocking. We recently celebrated the 10-year anniversary of the first commit to dbt Core. There are now over 100,000 teams using it in production every single week, and over a billion downloads. Many of the ideas dbt started with, like testing your data code and keeping data in version-controlled repos, were controversial a decade ago. Now they're defaults. ## dbt State: A caching layer that saves you money [dbt State](https://docs.getdbt.com/docs/deploy/dbt-state-about) is one of the most important features we've ever shipped. dbt State is a caching layer for dbt directed acyclic graphs (DAGs). If nothing has changed at the column level, the table level, or the code level, and you're inside the freshness window you've defined, dbt just skips that model. It doesn't need to build it. The funny thing about building a caching layer is that you don't know how much more efficient everything could be until you implement it. It turns out dbt State saves everybody a lot. That works out to a conservative 30%+ of a customer's dbt-driven compute bill. For us at dbt Labs, it was bigger, more like 64% of the compute bill. We've been able to cut about $400,000 annually from our budget for our underlying data platform. We added headcount as a result. The savings aren't just in production. Taylor shared that Fivetran had a large revenue model that took something like 40 minutes to run; once dbt State was turned on, it ran in about 40 seconds. The analyst team said they literally can't go back. They did see more queries as a result, which drives prices up slightly, but the experience was profoundly better. In development, dbt State makes your workflow really tight because you don't have to sit and wait to rebuild your entire development environment. In production, it saves you a tremendous amount of money. One thing to be clear about: dbt State is a paid feature, but it doesn't require the dbt platform. If you're using dbt Core plus [Airflow](https://airflow.apache.org/) today, you can still run dbt State, and we'd love it if you did. The lowest supported version is dbt v1.7 with a plugin; from v1.12 on, and certainly in v2.0, dbt State is an incorporated feature. ## dbt Wizard: A coding agent purpose-built for dbt There's one more feature we have that will supercharge your development times: [dbt Wizard](https://www.getdbt.com/product/dbt-wizard). dbt Wizard is a coding harness, similar to [Codex](https://openai.com/codex/) or [Claude Code](https://claude.com/product/claude-code), but tuned for dbt tasks. It turns out vertical-specific harnesses have a ton of benefits relative to generic harnesses built for all software engineering tasks. The difference is that we control the harness from the ground up. So we can tune it to perform better and be more token-efficient for the kinds of problems [analytics engineers](https://www.getdbt.com/blog/what-is-analytics-engineering) are solving all day, every day. dbt Wizard uses whatever model you run internally. You plug in your API key and pull from the same token budget, but it will frequently do a better and more efficient job than a generic agent. Try it out, and let us know in the [dbt Community Slack](https://www.getdbt.com/community/join-the-community) what your experience is. ## A foundation for the next decade A decade in, the ideas dbt introduced are now just how data work gets done. The next decade is about making that same trusted foundation work for agents as well as humans. dbt Core v2.0, dbt State, and dbt Wizard are the first big steps. For the full conversation, including the live Q&A, [watch the webinar recording](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). And if you want to see the speed for yourself, [create a free dbt account](https://www.getdbt.com/lp/dbt-free-account) and invoke dbt Core v2.0. That wow moment is the whole point. ## Frequently asked questions ### Will dbt Core continue to be open source? Yes. ### Are there plans to make dbt packages private? No. As Taylor put it, we think it's important that they remain open source and that anyone can access them. We'd still like you to run them on Fivetran plus dbt, but you don't have to. ### Are dbt Core v2.0 and the dbt Fusion engine the same thing? They're built on the same engine, and Fusion is a superset of Core. If you're already running dbt Core v2.0, there is no upgrade to Fusion; it's just a question of whether you want to flip on additional features. Many of them are free, so it's an exclusively better experience to run with them turned on. The [Fusion availability page](https://docs.getdbt.com/docs/fusion/fusion-availability) breaks down exactly what's included where. ### What advantages remain in Fusion if Core v2.0 uses the same engine? The core capability Fusion has that dbt Core doesn't ship is SQL parsing. Fusion natively understands the SQL you write, which powers improved editor capabilities and column-level lineage in your documentation, and gives us a foundation for future functionality like personally identifiable information (PII) classifiers that flow through the DAG. ### Is dbt Core v2.0 still a Python package, even though it's implemented in Rust? Yes. If dbt had started in the Rust community, it never would have been distributed via PyPI in the first place. But dbt grew out of the Python community, and there are over 100,000 deployments that start with pip install dbt. So, while much of the code is, in fact, Rust, we're continuing to make it available via PyPI. You can also install the binary directly. ### Is dbt moving to usage-based pricing? We do anticipate monetizing dbt State on usage. But dbt State is a brand new feature, and we often make independent decisions about how to monetize brand new things. There's no specific plan to change how the products and services that existed before dbt State are priced. Over time, we think of dbt more and more as infrastructure and want to monetize more in line with consumption. But that's a long-term arc, and we'll take every step carefully in consultation with customers. ### How do the dbt State costs actually work out? We haven't released final pricing yet. The entire point is that every dollar you spend on dbt State saves you more than that. As Taylor put it, we'll give you a dollar for 50 cents. With dbt State, you spend a lot less money on your compute provider, we charge you a little more, and your total cost of ownership goes down. When we deploy dbt State with a customer, we'll frequently run a proof of value in production for a couple of weeks. The before/after delta makes it a straightforward conversation. Over a hundred customers have adopted dbt State already. Everyone's environment is different, so it's worth testing to see how much you'd save. ### What's changing in the semantic layer? We released a new spec for the semantic layer in v2.0 and Fusion. Previously, you had to write a top-level semantic model block as a separate thing and map it to the model, which was clunky. Now, natively within the model YAML, you can declare a semantic model, mark dimensions in the same syntax you use to describe a column, and define metrics right in the YAML. We're also investing in making it more expressive, because folks coming from LookML sometimes hit patterns it can't support natively. That’ll take a good six to 12 months for the semantic layer, so keep an eye out for upcoming changes. ### If AI agents become the primary data consumers, do SQL files become obsolete? Some folks have said that in the future, the spec is the code, and everything else is an intermediate compilation artifact. There are versions of the world where that's true. But for the coming several years, we don't anticipate it. Intelligence is really expensive. Executing software is dramatically more resource-efficient than starting from natural language text every time. So we don't think SQL files are going anywhere, and neither are Python or Rust files. The ways we author them will change, and it's becoming easier to build the surrounding infrastructure of tests and documentation, but they aren't losing relevance. The previous data stack was about building analytics for humans. The emerging stack is about building context for agents. ontext engineering is a real, broad role, and analytics engineers will play a huge role in it. It's a tailwind for the career of every analytics engineer. ### Will Fivetran and the dbt platform become a single app? Today we haven't spent a ton of time on a single app experience, partially because the integration between the two products is already very tight. We'd rather invest the energy in innovation like dbt State, dbt v2.0, and dbt Wizard. Looking five years out, it's hard to imagine two separate auth systems for transformations and integrations, especially in a world where more of this gets built by agents. But our number one priority is integrating the teams and keeping up the velocity of shipping things that matter to customers. For the full conversation, including the live Q&A, [watch the webinar recording](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). And if you want to see the speed for yourself, [create a free dbt account](https://www.getdbt.com/lp/dbt-free-account) and invoke dbt Core v2.0. --- --- title: "Your next level starts here: A preview of dbt Summit sessions, by role" description: "Preview dbt Summit 2026 sessions by role: hands-on labs and breakouts for analytics engineers, data leaders, and execs." url: "https://www.getdbt.com/blog/dbt-summit-2026-sessions-by-role" date: "2026-07-20" authors: ["Daniel Poppy"] categories: ["Learn"] --- # Your next level starts here: A preview of dbt Summit sessions, by role Ten years ago, this community rewrote how the world works with data. Analytics engineers turned data teams into strategic drivers of the business. We're at another one of those moments. AI changes what it means to work with data. The consumers of your models are shifting from analysts in a BI tool to agents acting on their own, at machine speed and scale. Four out of five data leaders say their data isn't ready for enterprise AI. The teams who win this era are the ones who build the foundation everything else runs on. That's you. dbt Summit 2026 is where this community levels up together. September 15-18 at The Cosmopolitan in Las Vegas, with 100+ sessions across keynotes, breakouts, hands-on labs, and peer exchanges. We've pulled out just a sample of the sessions worth planning your week around and organized them by role, so you can jump straight to what fits your work. [Register here](https://www.getdbt.com/dbt-summit/registration). ## For analytics and data engineers You're the one shipping models. These sessions are about getting hands-on with what's new and hearing from peers who have already put agents to work. **Hands-on learning** [**Accelerating your deployment with dbt v2**](https://www.getdbt.com/dbt-summit/agenda/accelerating-your-deployment-with-dbt-v2). Get hands-on with the next-generation dbt engine and work through realistic analytics engineering problems in a dbt v2 environment. You'll leave able to predict how a model or source change will influence a run. **[Standardizing insights with the dbt Semantic Layer](https://www.getdbt.com/dbt-summit/sessions/standardizing-insights-with-the-dbt-semantic-layer).** Start from a dbt project, define semantic models and metrics, validate the logic and the grain, then confirm the metric holds up when it's consumed downstream. It's the difference between writing a definition and actually having one. [**Accelerating analytics with AI**](https://www.getdbt.com/dbt-summit/sessions/accelerating-analytics-with-ai). Implement a dbt feature end-to-end with AI assistance, reviewing it against a structured checklist for correctness, performance, and maintainability, then turning those standards into a reusable custom dbt agent skill. **[Migrate stored procs with dbt Wizard](https://www.getdbt.com/dbt-summit/sessions/migrate-stored-procs-with-dbt-wizard).** Everyone has the stored procedure nobody wants to touch. Turn a legacy proc into a dbt project you can maintain and test, using dbt Wizard to generate the model layers and `audit_helper` to validate parity so you can prove the migration is correct rather than hope. You'll finish by defining a Semantic Layer metric on top of the migrated models. [**Beyond the basics: Running dbt at scale on Microsoft Fabric**](https://www.getdbt.com/dbt-summit/agenda/beyond-the-basics-running-dbt-at-scale-on-microsoft-fabric). Build a production-grade medallion architecture on Fabric, orchestrated end-to-end through a dbt job. Incremental models, contract testing at layer boundaries, CI/CD, and metadata-driven orchestration with Fabric Data Pipelines so new sources onboard by config, not by hand. **Breakout sessions** **[YAML doesn't know why: Building the business context your agents are missing](https://www.getdbt.com/dbt-summit/sessions/yaml-doesnt-know-why-building-the-business-context-your-agents-are-missing).** The dbt MCP server, dbt skills, and dbt Semantic Layer give agents technical fluency. The context agents keep missing is organizational. Pedro Heyerdahl of Kilo Code shows how to build a living context layer that multiple agents can read from and contribute to. **[From AI experiment to production: How Okta governs context for agents at scale](https://www.getdbt.com/dbt-summit/sessions/from-ai-experiment-to-production-how-okta-governs-context-for-agents-at-scale).** Okta found that production AI depended less on a better model and more on a governed, discoverable semantic layer any agent could reason over from day one, built on dbt as the source of truth. Pooja Crahen shares how they got there. **[How to build governed agentic analytics on Amazon Redshift with the dbt Semantic Layer](https://www.getdbt.com/dbt-summit/agenda/how-to-build-governed-agentic-analytics-on-amazon-redshift-with-the-dbt-semantic-layer).** Amazon wires the dbt MCP server into the Amazon Redshift MCP Server, so an agent never authors a query, it resolves one compiled from a governed metric definition. The live demo traces one answer back to the dbt model that produced it and the tests that passed on it. **[No drift allowed: LangChain's context playbook with Hex](https://www.getdbt.com/dbt-summit/agenda/no-drift-allowed-langchains-context-playbook-with-hex).** Emily Hawkins and Logan Cochran of LangChain layered dbt, semantic models, and workspace guides into a context stack that turned their Hex agent into a source of truth the whole org trusts. They'll cover how endorsements guard against ungoverned data and how Context Studio closes the feedback loop. **[How Sigma manages the semantic layer with dbt and Dagster](https://www.getdbt.com/dbt-summit/agenda/how-sigma-manages-the-semantic-layer-with-dbt-and-dagster).** Matt Senick walks through the Dagster job that deploys every semantic asset at Sigma, Snowflake semantic views, Cortex Agents, Cortex Search Services, and Sigma Data Models, from a single dbt project each time something changes. **[Sweetwater's proactive data observability playbook with Datadog](https://www.getdbt.com/dbt-summit/agenda/sweetwaters-proactive-data-observability-playbook-with-datadog).** Daniel Gonzalez and Derek Andres on pairing dbt with Datadog to catch a long-running script delaying order updates before anyone noticed, and how dbt Mesh eased the shift to departmental ownership. **[SQL-first AI: bringing BigQuery AI into your dbt project](https://www.getdbt.com/dbt-summit/agenda/sql-first-ai-bringing-bigquery-ai-into-your-dbt-project).** Google's Alicia Williams and Jobin George on using AI.GENERATE, AI.CLASSIFY, and AI.SCORE to turn raw logs into structured insights without leaving your SQL workflow, plus the cost and latency tradeoffs of running AI at scale in dbt. **Peer exchanges** Peer exchanges are small-group, discussion-first sessions. You bring your experience and your notepad. **[Agents, MCPs, and buzzword fatigue: What AI actually changes for analytics engineers](https://www.getdbt.com/dbt-summit/sessions/agents-mcps-and-buzzword-fatigue-what-ai-actually-changes-for-analytics-engineers).** New AI tooling launches weekly and the terminology multiplies faster than the problems it solves. XiaoHan Li of Xebia hosts a hype-free conversation about which tools actually stuck, who owns the logic when AI writes your models, and the skills worth investing in as more of the boilerplate gets automated. **[How to build a successful data career](https://www.getdbt.com/dbt-summit/sessions/how-to-build-a-successful-data-career).** Three practitioners, Millie Symns of Justworks, Silja Märdla of Bolt, and Bruno Lima of phData, trade practical patterns for building career momentum as AI shifts what's expected of the role. Expect honest talk about durable skills, cross-industry moves, and making your impact visible without defaulting to "just become a manager." ## For data team leaders You're deciding how your team scales, standardizes, and stays ahead. These sessions are about rollout patterns, governance, and positioning your team for what's next. **Hands-on labs** **[Scaling trusted self-service for dbt stakeholders](https://www.getdbt.com/dbt-summit/sessions/scaling-trusted-self-service-for-dbt-stakeholders).** Scale dbt beyond the build team by helping stakeholders find, understand, and reuse trusted data products without turning everyone into a developer. Walk you through documentation patterns, ownership, and a stakeholder access model that expands governed consumption while protecting your development workflow. **Breakout sessions** **[From selection to scale: How ING is operationalizing dbt across a global bank](https://www.getdbt.com/dbt-summit/sessions/from-selection-to-scale-how-ing-is-operationalizing-dbt-across-a-global-bank).** Jarno Boeijink shares how ING drives governed enterprise adoption inside a regulated bank, and how the dbt Semantic Layer and the dbt MCP server are opening new ground for natural-language analytics and AI-assisted development. **[Governed by default: How data teams at Nordstrom turn dbt governance into an AI advantage](https://www.getdbt.com/dbt-summit/sessions/governed-by-default-how-data-teams-at-nordstrom-turning-dbt-governance-into-an-ai-advantage).** Nadine Bruxel makes the case for dbt as the control surface for safe AI: freshness as an AI SLA, tests and contracts and lineage as guardrails, and a conversational agent built on top of all of it. Governance-first is how Nordstrom gets to AI readiness. **[Multi-agent dbt orchestration at Riot Games: Redefining the analytics engineering SDLC](https://www.getdbt.com/dbt-summit/sessions/multi-agent-dbt-orchestration-at-riot-games-redefining-the-analytics-engineering-sdlc).** Jessica Zhang shows how Riot Games safely coordinates multiple agents to read metadata, translate legacy logic, and generate pull requests, all behind read-only guardrails. If you want to know how far multi-agent workflows can go in production, this is the session. **[Scaling dbt on Amazon Redshift: how KOHO cut transformation runtime 70% without rewriting a single model](https://www.getdbt.com/dbt-summit/agenda/scaling-dbt-on-amazon-redshift-how-koho-cut-transformation-runtime-70percent-without-rewriting-a-single-model).** KOHO re-architected from a single Redshift cluster to a Hub and Spoke model and a data mesh, cutting its nightly dbt runtime from 6.5 hours to 2 in a config-only migration, no model rewrites, no downtime. **[From Data to AI: Why Microsoft Fabric and dbt are better together](https://www.getdbt.com/dbt-summit/agenda/from-data-to-ai-why-microsoft-fabric-and-dbt-are-better-together).** Roy Hasson, Pradeep Srikakolapu, and Abhishek Narain of Microsoft on how OneLake, unified security, and AI-powered experiences pair with dbt's testing, lineage, and documentation to cut fragmentation and build a trusted foundation for analytics, AI, and agentic workloads. **[From prototype to production: How DoorDash built a scalable analytics SDLC with dbt and ThoughtSpot](https://www.getdbt.com/dbt-summit/agenda/from-prototype-to-production-how-doordash-built-a-scalable-analytics-sdlc-with-dbt-and-thoughtspot).** Harsha Reddy on leading the modernization of DoorDash's analytics stack onto dbt and ThoughtSpot, consolidating a large, multi-team org onto a platform that now serves more than 10,000 users. **Peer exchanges** **[Beyond the bottleneck: Position your analytics engineering team as a strategic force](https://www.getdbt.com/dbt-summit/sessions/beyond-the-bottleneck-position-your-analytics-engineering-team-as-a-strategic-force).** Kasey Mazza of HubSpot leads a discussion on moving your analytics engineering team from a service desk to a strategic driver, with the framing and language to make that shift stick with leadership. **[How to build a successful data career](https://www.getdbt.com/dbt-summit/sessions/how-to-build-a-successful-data-career).** Worth the crossover for leaders too. Millie Symns, Silja Märdla, and Bruno Lima swap patterns on durable skills and career growth, useful for anyone coaching a team through the AI shift. ## For business leaders and executives You're weighing where data investment turns into measurable outcomes. These sessions lead with results, governance, and the business case for a strong data foundation. **Breakout sessions** **[Real-time analytics at Bilt: Architecture and approach](https://www.getdbt.com/dbt-summit/sessions/real-time-analytics-at-bilt-architecture-and-approach).** James Dorado and the Bilt team walk through the architecture behind their real-time analytics, powering audience targeting and offer execution on fresh data. **[An AlphaSense case study: Scaling AI on enterprise data with dbt-first governance and context from Euno](https://www.getdbt.com/dbt-summit/sessions/an-alphasense-case-study-scaling-ai-on-enterprise-data-with-dbt-first-governance-and-context-from-euno).** Sarah Levy of Euno and Brad Levy of AlphaSense show how dbt-first governance, paired with automated context, keeps AI decisions explainable and traceable back to governed source data. **[Governed by default: How data teams at Nordstrom turn dbt governance into an AI advantage](https://www.getdbt.com/dbt-summit/sessions/governed-by-default-how-data-teams-at-nordstrom-turning-dbt-governance-into-an-ai-advantage).** The executive read on the Nordstrom story: governance is the thing that makes AI on your data safe to trust and safe to scale. Nadine Bruxel shows how a governance-first foundation turns into a real advantage in the AI era. **Peer exchanges** **[Empowering stakeholders in the age of AI](https://www.getdbt.com/dbt-summit/sessions/empowering-stakeholders-in-the-age-of-ai).** Lexi Galantino of Zipline hosts a conversation on what "talk to your data" actually takes: which models make it work, how you keep the answers correct, and what the role of the data team becomes when stakeholders can propose their own changes. ## Build your week around it These are a fraction of the 100+ sessions on the agenda. Anchor your schedule around the two keynotes, Level Up on Wednesday and the Community Keynote on Thursday, then fill in the breakouts, labs, and peer exchanges that map to what you're building next. Registration is open now, and the $1,695 registration includes a free training and certification while spots last. Bring home the templates, playbooks, and patterns you can put to work immediately. Your next level starts here. [Register for dbt Summit](https://www.getdbt.com/dbt-summit/registration). --- --- title: "OSI is now Apache Ossie (Incubating)" description: "Apache Ossie is currently undergoing incubation at The Apache Software Foundation (ASF)." url: "https://www.getdbt.com/blog/osi-is-now-apache-ossie" date: "2026-07-13" authors: ["Quigley Malcolm"] categories: ["Partnerships"] --- # OSI is now Apache Ossie (Incubating) If you've been following the Open Semantic Interchange (OSI) project, the open specification for semantic layer and ontology, there's an important update. The project has been accepted into the Apache Incubator. Along with this transition the name is changing to Apache Ossie (Incubating). The spec, the community, and the mission haven't changed, but the name, governance home, and long-term trajectory have. ## Why the new name? When work on this initiative was started, it was called the Open Semantic Interchange. This quickly got shortened to OSI, which became the GitHub repo name. Unfortunately this has caused some confusion along the way as the acronym OSI is used frequently to refer to Open Source Initiative. Now although it might be fun to say OSI OSI (Open Source Initiative Open Semantic Interchange), the community decided it was best to rename the project. Through discussion the community decided on Ossie, and with acceptance into the Apache incubator, it is now Apache Ossie (Incubating). In addition to the rename, a mascot has been chosen, a kangaroo. The Ossie Kangaroo is dedicated to carrying semantic metadata in its pouch from platform to platform. That is, it’s making your data hop. In short: - The project is Apache Ossie (Incubating) - Any reference to "OSI" in the project are historical (and will slowly be removed) - There is a kangaroo logo If you've been building on Open Semantic Interchange, nothing breaks. The name changed, but the spec didn't. ## What is Ossie? Ossie is an open specification for both semantic layer and ontology. It defines a vendor-neutral format for expressing business metrics, dimensions, relationships, as well as broader business concepts and rules. It allows any tool or platform in your semantic layer stack to produce and consume semantic definitions without loss of meaning. The problem it solves is important: it ensures that a given business concept (say, "Monthly Active Users") can be defined, interpreted, and resolved consistently across an organization's CRM, data warehouse, and BI tools. When a human analyst or an AI agent runs a query, they shouldn't have to guess which definition is correct. Ossie provides the shared, machine-readable format that encodes not just the data but the intent and business meaning behind it. ## Why the Apache Software Foundation (ASF)? Incubating Ossie with the Apache Software Foundation ensures that it remains an open standard with no single controlling entity. The goal of Ossie is to provide industry-wide standardization of semantic data, and to that end ensuring that it has a vendor neutral ground to operate in is imperative. Under incubation, Ossie operates with public mailing lists, GitHub-based development, a formal discussion-and-vote process for spec changes, and committership earned through contribution rather than employer affiliation. Note that as part of this transition, all mailing lists referring to Open Semantic Interchange will be retired; community members should use the ASF-provided project resources that are linked below instead. ## Importance of the Ossie community Ossie didn't start as a single-company project, it has been a community effort. Since the repository opened in November 2025: - More than 100 commits and 35 merged pull requests have landed from contributors at Snowflake, Salesforce, Databricks, dbt Labs, RelationalAI, GoodData, and Honeydew - The participating coalition has grown from 17 launch partners to [more than 50 organizations](https://www.snowflake.com/en/blog/open-semantic-interchanges-specs-finalized/) - Three working groups (Metric Language, Catalog, and Ontology) operate with dedicated leads, meetings and public channels - Implementations including the Ossie-to-dbt Semantic Layer converters and an Apache Polaris™ converter are already merged ## What's next dbt Labs was one of the founding organizations behind Open Semantic Interchange, and we'll continue as an active contributor to the project as it grows under ASF governance. As with any Apache project, the community will decide the direction together. That said, there are a few areas we're excited about and hope to work with the community to contribute proposals for: - Deepening the spec's expressiveness to accommodate what real enterprise models demand, including an expression language spec, advanced metric logic, windowing functions and complex relationships - Building converters for additional platforms and frameworks so that adopting Ossie doesn't require ripping out what you already have - A standardized semantic query specification that any engine can support - Integration with Apache Polaris so that semantic models are discoverable directly from the catalog None of this is predetermined. It will go through the same open discussion-and-vote process as everything else in the project. ## Get involved Ossie is transitioning to ASF infrastructure as part of incubation. Watch for updates on the new [project website](http://ossie.apache.org), join the [development mailing list](mailto:dev@ossie.apache.org), collaborate on [GitHub](https://github.com/apache/ossie) and join the [Ossie Slack workspace](https://join.slack.com/t/apache-ossie/shared_invite/zt-42i1xkgy8-7YQtKEDq7v~mceFmdiLhkA). Whether you're building an AI agent, BI tool, or a query engine that needs to understand business context, Ossie is the community working to make sure you don't have to tackle semantic interoperability alone. --- --- title: "The productivity gains hiding in your data infrastructure" description: "Budgets aren't growing, but the work is. See how dbt customers recouped 58.7 FTEs in capacity, worth $1.75M a year." url: "https://www.getdbt.com/blog/data-infrastructure-productivity-gains" date: "2026-07-08" authors: ["Daniel Poppy"] categories: ["Product"] --- # The productivity gains hiding in your data infrastructure Demand for data is exploding thanks to AI. Some experts estimate that global spending on data centers capable of handling advanced AI workloads [could hit $7T by 2030](https://www.mckinsey.com/industries/technology-media-and-telecommunications/our-insights/the-7-trillion-dollar-data-center-build-out-how-industrials-can-capture-their-share). Data teams aren't necessarily getting more resources to deal with it, though. When we talked to companies, we found that only 36% of data teams reported increasing budgets. That means data teams have to shoulder more work with existing capacity. AI itself helps, of course. Tools like [data copilots](https://docs.getdbt.com/docs/dbt-ai/copilot-overview) reduce the burden of generating code that produces clean, high-quality data for AI agents. But AI can only do so much. The hard problems still need humans to solve them. And currently, those humans are mired in maintenance, working assiduously to keep the data house of cards from falling down around them. The good news is that data teams can free up significant capacity by taking a modular, reusable, automated approach to managing AI and analytics data workloads. This isn't just speculation. [A recent IDC report quantifies](https://www.getdbt.com/resources/the-business-value-of-dbt-idc-report) exactly how much using a platform like dbt can save at every step of the data lifecycle. ## The maintenance nightmare No data team handles data from a single source. Everyone is constantly wrangling data from a variety of data storage platforms and formats. The advent of AI has made this even more of a challenge. Data teams aren't just dealing with structured relational data and semi-structured data sources any longer. They're also mining PDFs, emails, and social media posts for insights. All this means that most teams end up taking a scattershot approach to managing data pipelines. Most are written on the fly as quick and dirty one-offs meant to get the job done. The result? A maintenance nightmare. - Pipelines are brittle and prone to breaking. Respondents in our annual [State of Analytics Engineering Report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) reported that they spend a significant amount of their time maintaining data sets, platforms, and infrastructure. - Most work isn't reusable across data pipelines, forcing teams to rebuild what they need from scratch each time. - Testing, deployment, and review are slow, manual processes, if they exist at all. - New data contributors face a long ramp-up time learning how to navigate heterogeneous data systems. Many systems end up being too complex for non-technical contributors to use efficiently. Each of these factors eats up precious time that data team members could instead be spending on more strategic work, such as streamlining data intake, optimizing storage, improving overall scalability, and automating key processes for faster delivery and more consistent quality. ## dbt: Unlocking capacity without hiring For years, dbt has served as the [data control plane](https://www.getdbt.com/blog/data-control-plane-why) for companies worldwide. By taking a single, vendor-agnostic approach to modeling data, testing changes, and orchestrating data pipelines, dbt reduces complexity, boosts reusability, and improves quality across all analytics and AI data workloads. The results, as summarized by the [IDC Business Value report](https://www.getdbt.com/resources/the-business-value-of-dbt-idc-report), are real and measurable. IDC interviewed eight enterprise companies that use dbt, with an average of 20,000 employees and $16.75B in annual revenue. In headcount numbers, across various data functions, companies reported recouping the equivalent of **58.7 FTEs' worth of capacity**. The businesses reported doing more with less headcount across all data functions: - Analytics and data teams: +25 FTEs - Developers: +11.2 - Business analysts: +17.6 - Data governance: +2.5 - Platform management: +2.4 The analytics/data team savings alone are the equivalent of $1.75M of additional business value delivered every year just by using dbt. In each case, functionality provided out of the box by dbt enabled teams to achieve these gains: ## A speed boost across the data stack These savings come from reducing the cycles required to perform basic data tasks. At every step of the data lifecycle, dbt enables teams to deliver more in less time: - Report delivery drops from 16.3 days to 8.4 days, a **49%** acceleration - Testing is **44% faster** for new apps, and **46% faster** for pipeline updates - Development cycles are **41%** and **37% faster** for new features and updates, respectively - Teams reported delivering new solutions to market **34%** faster, and scaling **33%** faster "dbt has significantly improved developer collaboration and increased our development velocity," one customer told IDC. "It introduced a structured deployment process through its integration with Git." ### Faster onboarding and reuse Teams also reported significantly faster onboarding times. Onboarding dropped from an average of 3.3 weeks to 1.7 weeks, a **47% acceleration**. The driver? The availability of existing models. With data modeled consistently and available for self-service discovery via [dbt Catalog](https://www.getdbt.com/product/dbt-catalog), new hires have access on average to 101,970 reusable models from day one. This means the productivity gains offered by dbt aren't a one-time bump. Each new model is an investment that compounds into the future as new hires plug existing data transformation code into their own pipelines. "dbt platform is very easy to use for new employees," reported one customer. "It's all templates and SQL, so people are comfortable with it." ### Reducing the data quality tax Testing, automated deployment, and clear documentation have a measurable impact on data quality. Companies reported **33% fewer data quality issues** due to dbt. They also reported **35% fewer late-data** and **13% fewer incomplete-data** instances. This is important because data teams historically spend much of their time responding to and fixing data quality issues retroactively, after they make it to production. Studies in software engineering have shown that fixing bugs in production is significantly more time-consuming and expensive than ensuring those defects never ship in the first place. A bug that costs $100 to fix in the design stage of a project [may become a $10,000 problem](https://www.forbes.com/councils/forbestechcouncil/2023/12/26/costly-code-the-price-of-software-errors/) if it makes it to production. "We use dbt to run tests such as checking for duplicates, missing values, and incorrect data," another customer said. "dbt helps identify issues and traces them back to their source, allowing us to fix problems at the system level and deliver more reliable data." ## The capacity was never missing None of these gains come from a single feature. They come from applying a single discipline, the [Analytics Development Lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle), to every stage of data work: - **Build**: Reusable models replace one-off scripts. Teams solve each problem once, not repeatedly. - **Test**: Automated checks catch errors before they ship, not after they've reached a dashboard or an AI model and led to a costly business decision. - **Deploy**: Continuous Integration and Continuous Delivery (CI/CD) move changes automatically through a rigorous change management process, not via manual pushes and after-the-fact firefighting. - **Operate and discover**: Data lineage makes it easier for engineers to identify the root cause of data issues. Data cataloging and health reports enable self-service so that everyone can leverage what already exists instead of re-inventing the wheel. This capacity wasn't missing. It was buried in maintenance and the errors caused by manual work. dbt frees up this capacity so that teams can ship more and ship faster using the resources they already have. To learn more about how dbt saves companies time and money, [download the full IDC report](https://www.getdbt.com/resources/the-business-value-of-dbt-idc-report). To get started enhancing your data productivity, [talk to us today](https://www.getdbt.com/contact) about how dbt fits into your existing data stack. --- --- title: "Solving dashboard errors in minutes: How Integral Ad Science used MCP to connect agents to dbt and Databricks" description: "Integral Ad Science used MCP to connect AI agents to dbt and Databricks, turning hours of dashboard debugging into minutes." url: "https://www.getdbt.com/blog/mcp-dbt-databricks" date: "2026-07-07" authors: ["Daniel Poppy"] categories: ["Product"] --- # Solving dashboard errors in minutes: How Integral Ad Science used MCP to connect agents to dbt and Databricks Stop us if you've heard this one. A user goes into a report on a BI dashboard and starts asking questions of the embedded chatbot. They're drilling, filtering, checking the numbers… And all of a sudden, things don't add up. This isn't an uncommon problem for analysts in the AI data age. Mars Dauer, Senior Director of Enterprise Data and AI at Integral Ad Science (IAS), and his team ran into the same issue. They knew they had the information that their AI agent needed; it was already there in dbt and Databricks. The challenge was how to supply it to the agent so that it could solve what Dauer and his team call the "BI why" problem. At Databricks Data + AI Summit 2026, Dauer shared how teams can supply this missing context via model context protocol (MCP) servers that connect AI chatbot agents to dbt and Databricks. In this article, we'll dig into the architecture he used to solve the problem, the decisions made, and the honest lessons about what worked and what the team learned. ## The BI "why" problem: What copilots can't answer BI chatbots have what Dauer calls a "BI why" problem. They can tell you what the numbers say. However, they can't explain the context and reasoning behind them. Let's assume, for example, that the company is a supermarket chain. An analyst is digging into a revenue report and breaks it down by category — and Food jumps out as unreasonably high. But why? Can the AI agent tell them where the disconnect is? This is where the analyst realizes (to no one's surprise) that the agent can't. It can summarize the chart; it can tell them what filters are applied. But it has no real understanding of the business logic or the transformations that produced those numbers. So, the analyst digs deeper. They open [Looker](https://cloud.google.com/looker) and check the LookML. That points to a dbt model, so they jump into the model and read the SQL. That model leads to three upstream marts, so they open more SQL in more tabs. They eventually run a few validation queries in [Databricks](https://databricks.com) and discover that the models have miscategorized a popular drink as food. Looking at the architecture of a typical AI chatbot system in BI tooling, it's easy to see why the bot couldn't unearth this issue. The BI copilot can see only the topmost layer of the data stack: - The BI tooling itself - The semantic layer that defines a unified layer for key metrics and governance - The mart models containing the gold standard data [Slide 5] However, it can't see everything _under_ that—in this case, an errant CASE statement. The agent can't see the business logic built into the dbt intermediate models, or the sources in Databricks it could use to validate and correct this error. ## MCP: Your USB-C for LLM tools Dauer knew that data was available, which meant the answer was somewhere. The issue was supplying it to the BI copilot so that it could solve the issue in minutes, not hours. The answer his team hit on: MCP. [MCP defines a standardized framework](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) that enables AI applications, such as agents, to talk to tools and understand their capabilities without needing to know their internal APIs. Dauer refers to MCP as the USB-C for large language models (LLMs): the agent has one kind of port, and any tool that exposes an MCP server with the right shape can plug into it. Before MCP, every integration meant a new bespoke client. With MCP, tools advertise their capabilities, and any MCP-compatible agent can discover and utilize them. For the project, Dauer's team needed three MCP servers to solve the issues they'd detected: - The **dbt MCP server** (the structured context layer), which connected the agent to tested, version-controlled, and documented model definitions — including model metadata, [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), SQL, and field/table descriptions; - A **Databricks SQL MCP server** (governed data validation layer), which enabled read-only query execution and [Unity Catalog](https://www.databricks.com/product/unity-catalog) governance; and - A **Looker MCP server** (the BI on-ramp), providing access to SQL parsing for [explore URLs](https://docs.cloud.google.com/looker/docs/viewing-and-interacting-with-explores) and dashboard tiles. The team had been using all these MCPs for a while in coding agents like [Cursor](https://cursor.com/), [Claude Code](https://www.anthropic.com/product/claude-code), and [Visual Studio Code](https://code.visualstudio.com/). Using them in conversational/analytics agents felt like the logical next step. ## How the analytics agent works The agent has five core capabilities. Dauer thinks about them in three buckets: **Lineage and logic.** Given a column or metric, it traces that object from the mart back to the source. It reads the compiled SQL at each hop, and explains in plain language, what each transformation is doing. **Validation.** Once it has the lineage, the agent can validate against the actual data in Databricks using read-only SQL. It can also perform freshness checks and monitor pipeline health. **Impact analysis. **If you change a column upstream, the agent tells you what breaks downstream, and because it reads the logic in each downstream model, it can reason about whether your change matters semantically, not just structurally. The agent has two personas depending on who's using it. In **analyst mode**, it returns the answers that data producers and analysts need: full technical details, including model names, compiled SQL, column catalogs, directed acyclic graph (DAG)-aware lineage, and raw validation query results. In **business mode**, it gives only the information that matters to decision-makers: plain English answers (without jargon or model names), business context, and summarized findings. As Dauer emphasizes: "One agent, two voices: analyst mode shows model names and compiled SQL; business mode strips both for plain English. Same investigation, different answer for a CMO vs a data engineer." ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/09d66b3d44fb7c377a210cc235d98eeb9c2c3d72-2046x1068.jpg) For Integral Ad Science, Looker was the primary point of entry. The agent was in the BI layer so that users didn't have to context-switch. The agent runs as a window to the right in the main Explore to provide the best user experience. The backend is a FastAPI server running on Cloud Run. It handles auth through Google OAuth and an authorized group, manages sessions, and streams responses back over Server-Sent Events. Then there's the agentic runtime, which uses Google's Agent Development Kit. It's composed of a single root agent and four sub-agents. The root agent takes the user's question, decides which specialists to call, fans out the work, collects the results, and formats the final answer. The LLM layer here is deliberately provider-agnostic via the LiteLLM library. This means each subagent can target Databricks Foundation Model APIs, [Gemini Enterprise Agent Platform](https://cloud.google.com/products/gemini-enterprise-agent-platform), Anthropic, [OpenAI](https://openai.com/) - whatever best fits the job. This means subagents can use smaller, less expensive models for simpler jobs, such as routing or analyzing metadata. Finally, the MCP layer connects to dbt platform, Databricks SQL, Looker, and a few internal tools. All of the company's source systems are dbt projects built upon Unity Catalog in Databricks. ## Why MCP, and other design decisions A key question is: why MCP? Why build this using a universal adapter as opposed to creating custom integrations? A few reasons: **Speed.** Adding a new tool source meant pointing at a new MCP server, not writing a custom API client. This enabled dbt and Databricks to come online in days, not weeks. That speed compounds when you're still experimenting. **Decoupling.** Using MCP means the agent doesn't need to know how dbt, Databricks, or Looker work. Dauer's team could swap out any of these tools, and the agent's code wouldn't have to change. **Governance.** This is the one that matters most to security. The MCP server controls what's exposed. Integral can scope tool access, audit at the server boundary, and avoid handing raw API credentials directly to the agent. Databricks provides governed query execution; Looker provides the BI on-ramp. dbt provides the semantic and lineage layer that makes the whole investigation possible. It's the reason the agent can trace a metric from a dashboard tile all the way back to a source table — and trust what it finds there. Why rely on dbt MCP for lineage? Because dbt treats lineage as a first-class concept. The agent doesn't reconstruct the graph from scratch or hope documentation keeps pace with reality — it reads dbt's tested, version-controlled graph of how data actually flows. Alongside the graph, the agent reads the compiled SQL itself — and that's the ground truth. LookML is a model of a model. Documentation is what people intended (or hoped) the model did. Compiled SQL is what actually runs in the warehouse. The lineage tells the agent where to look; the SQL tells it what's happening there. ### Trade-offs and lessons learned Every architectural choice was a trade-off. Some of the key trade-offs that Dauer's team made: **Multi-agent vs. single-agent** The team started with a single agent. As tool complexity grew and prompts got longer, they decided to go multi-agent — specialists under a main routing agent. They landed at four. Every additional sub-agent adds a routing failure mode. The orchestrator has to choose the right specialist, and when it chooses wrong, they wouldn't always get a clear error. That led Dauer to make a rule: decompose only when the specialist has a genuinely different job and needs a meaningfully different system prompt. If two agents both query Databricks, and one is called "validation" while the other is called "exploration," they probably should be one agent. **Embedded vs. standalone chat** A standalone chat is tempting because you can ship it quickly. But it forces the user to copy context, meaning that every investigation starts with thirty seconds of context-pasting. The embedded Looker extension avoids that: it captures the dashboard/explore context automatically, so the question and the data live on the same surface. The trade-off here is real. Supporting another BI tool means redoing some of this work each time. But the savings per investigation compound. Bottom line: if your BI tool has an extension SDK, use it. **Per-sub-agent model tiering vs. one model for everything** The easy default is to pick one frontier model, use it for every sub-agent, and move on. IAS doesn't, Dauer said. This means simple agent tasks, such as the Looker resolver, aren't forced to use the most expensive model. They can save the more expensive model for operations such as the dbt analyst, where lineage reasoning over compiled SQL greatly benefits from the extra firepower. The rule: match the model to the work. The user doesn't feel a quality hit because the critical reasoning step still gets the strongest model. ## Only scratching the surface What used to take an analyst an afternoon, Dauer said, now takes minutes. And yet he feels like his team is only scratching the surface. The dbt platform MCP server exposes ~37 tools. IAS's solution only uses seven of them. That's actually a testament to how teams can get started narrow and expand without rebuilding anything — the breadth of the MCP surface area means the integration grows with your use case. Some of the tools the team is looking forward to including are: - Text-to-SQL with project context to generate SQL that knows the project's model names, joins, and conventions. - Code generation to create staging model YAML and sources. Leveraging this, agents could not only investigate issues but scaffold a fix and submit it for human review. - [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage?version=2.0&name=Fusion) via the Fusion engine — the agent already traces column-level lineage today by reading compiled SQL and inferring how columns flow. The Fusion-powered tool would give the agent that same lineage as a deterministic graph straight from dbt's compiler — replacing LLM inference with explicit column-to-column edges. Faster, cheaper, more precise. "The most interesting agents in this space probably haven't been built yet," Dauer said. "And, honestly, that's the part I'm most excited about." --- --- title: "A guide to implementing AI data pipelines" description: "AI coding took off. AI pipeline management didn't. Here's a practical playbook for closing that gap." url: "https://www.getdbt.com/blog/a-guide-to-implementing-ai-data-pipelines" date: "2026-07-06" authors: ["Stephen Thibeault"] categories: ["Insights"] --- # A guide to implementing AI data pipelines There’s a startling pair of statistics in dbt’s newly released [2026 State of Analytics Engineering](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) report: 72% of teams prioritize AI-assisted agentic coding for data work, while only 24% prioritize AI-assisted data pipeline management. Here are my thoughts on why this is happening and, if this is the case in your own org, how to go about building AI into your everyday workflows. ## Why AI pipeline management is lagging behind AI coding Looking at these numbers in relation to each other really shows the difference between **AI adoption** in terms of giving developers AI tools to work with, and **AI integration** by working AI into the data infrastructure itself. AI-assisted coding is very self-driven. You can be on your individual system, working on your own, using an AI coding tool on whatever dbt project you personally have pulled down. But using AI for pipeline management, like looking for job errors and feeding that data into an agent, has to happen at the team level. I can absolutely spin up an agent in a five-minute Claude Code session, but actually building the bots that would monitor a data pipeline takes time, dedication, and collaboration from basically the whole team. One of the largest clients I'm currently working with has a team of 80 engineers throughout their entire company that touch their dbt project. But when it comes to actually orchestrating that project, it's a team of maybe five engineers managing the orchestration out of the dbt platform. I think that's why we're seeing a slower uptake on using AI to assist data pipeline management, because it has to go through different approvals and the team has to be aligned on what exactly they want that to look like. Ultimately, the difference in those numbers are really about accessibility: AI coding is individual, personal, and easy for developers to adopt, versus actually building and managing AI pipelines in production as a team. ## Your data stack is already AI-ready The typical data architecture for many orgs today follows the ELT pattern: data is extracted from source systems, loaded into a data warehouse, and transformed for downstream consumption by business users. **A lot of bringing AI into your daily workflows does not necessarily change that stack. **You’re just layering AI over it, in the same way that dbt is currently layered over your data warehouse. The original data stack structure is still very solid, so we're just keeping that; AI now exists on top and covers the whole thing (or, alternatively, most of the original stack, because you absolutely can add this AI layer in pieces). ### Where to start The first step to implementing this new AI layer is to start with the highest-value layers. Here are places to consider adding in AI. **Code review** is where you can get the most out of AI by integrating it for things like reviewing PRs, CI checks, flagging issues like _'Hey, this is missing. We need to take a look at it.' _This is a natural starting point for introducing AI into data pipelines and workflows because it’s where developers are already using AI for their individual work. **Error triage** is a logical place to layer in AI: making AI part of your code review process, using it for error tracking in your data transformations, and even in your data loading workflows. In practice this would look like setting up AI to do that initial triage work and then setting up alerting, be it through email, Slack, PagerDuty, or whatever system you use for error communications. **Extract/load** is also a small but very targeted opportunity. Here, AI could be useful for evaluating what's coming in, making sure nothing coming in is messed up, evaluating errors, etc. Your** ticketing intake system** can be another great place to start. One dbt customer created a ticket intake chatbot to make sure they had access to all relevant context around requests. This chat interface gathered information about the problem that they're trying to solve, rather than simply intaking a verbatim request that essentially dictates a solution. Business users often don't know what the data team has access to or even really grasp the whole breadth of what that team can do. Applying AI here to create collaboration can produce solutions that give more than the person requesting it even realizes is possible, which transforms a data team from order takers into internal consultants. **Ultimately, each layer of the stack from intake to data loading to data warehousing to consumption have use cases for AI. **It's all about identifying those use cases and implementing them into that workflow, and you find them by looking for places where the addition of an AI tool would be helpful for you and your team overall. ## A playbook for getting started For an applied example, let’s say that you are currently the leader of a small data team of 10 people and you provide analytics to the sales team. You have a current modern data stack that consists of five ingestions, Snowflake as your data warehouse, [dbt to do the transformations](https://www.getdbt.com/blog/data-transformation), and Tableau for your data visualizations. A typical day-to-day for this team right now looks like this: business users put requests in, your data teams take those requests and create what's needed for the business, and finish with going through the review process. Where does AI enter into this picture? ### Step 1: Find the first, low-stakes pain point The first step in determining what you should do with AI is finding your use case. You do this by asking,_ 'Hey, where are our biggest pain points currently? Do we not have enough time to do extremely thorough code reviews? Or do we get incredibly vague information from business requests? Do we frequently get failed loads from our loading tool into the database due to changing APIs or other reasons?’_ **Code review **is the example we will use here as it is what a lot of people start with to really get their feet under them because, if something goes wrong, there is typically already an assigned human in the loop. To do this, you would: 1. Start by deciding orchestration, figuring out what automations exist and how you want to run those automations. Are you going to just use something like GitHub Actions? Do you have an external orchestrator that will run this process whenever a pull request is created in your change control tool of choice? 2. Next, determine which AI company you want to use (which is often determined by whichever one(s) your company already has enterprise agreements with) and which of their AI models you want to use. Obviously, because it's your code, you want to make sure that there are agreements for security, for not using your data for training, etc. All this should be handled in tandem with the IT team and the security team. 3. Build the behavior as skills: You would put together a skill that says, 'Hey, you're a code reviewer. These are our coding standards. We don't use leading commas. We always use CTEs in this specific format, we're using the dbt standards.' Whatever is important to you, it all goes here. _Pro tip: dbt has a [curated collection of agent skills](https://github.com/dbt-labs/dbt-agent-skills) that help AI agents understand and execute dbt workflows more effectively._ 4. Add the business context into the skill: You would also put information about the business, like the sales team expects this certain format, here's the basic structure of our data. You would build this all in markdown, use it as a skill that whatever AI tool you selected to automate this process can actually follow those instructions. ### Step 2: Set up a pilot Now, it’s pilot time for your AI-powered automated code review: roll it out, then let it run, monitor it closely, make sure you're still doing human code reviews. You’ll want to: - **Test and iterate before going wide**. It all comes down to testing that automation. So either putting it on a small test repo or putting it out there, but only running it kind of ad hoc to make sure that it works. So, doing a couple of code reviews, giving it some code, refining what that skill looks like, refining what you're sending to the AI. - **Pilot small and expand outward.** Your pilot is what builds that trust over time. Just releasing it to 500 people all at once will just create a mess trying to keep up with managing it. - The final goal is up to you, but for most teams this will be putting the AI code review agent on your main repo so that whenever somebody pushes to that repo, they're going to get this automated code review. ### Step 3: Gather feedback, create evals Let your people who are excited do the actual work of setting the pilot itself. When it comes to evaluation, however, you need to have people from multiple walks of life do that review so that you can make sure you're getting a well-rounded understanding of what's working and what's not working. The best feedback is going to come from a standard senior member on your team that is perhaps indifferent or maybe even a little combative when it comes to AI adoption. They'll be able to spot things that someone who's very excited about it may not spot. Have more senior members of the team who do a lot of code reviews also review the output and see if they can spot things that the AI is not doing correctly. This feedback can then be used to refine the Skill you are utilizing. Once those more senior members of the team feel comfortable with the results that you’re getting, now you work with them to create the eval for keeping things accurate over time (we’ll talk more about evals shortly). Finally: remember that this is all brand new, so don't feel like your pilot has to be perfect out of the gate. Making mistakes is how you learn not to make those mistakes; that's the IT way of life. The rite of passage for every data engineer out there is accidentally dropping a production table or breaking something in a transaction. But I predict that pretty soon another rite of passage for data engineers will be having an AI pilot that just didn't work out. Because, after all, how can you truly understand something unless you've broken it and then had to fix what you broke? ## Maintenance and evals are part of the planning process The crucial work in integrating AI into your data pipelines isn't actually building the pipelines, but making sure they stay reliable once they’re up and running. The most important message I can give you is that **integrating AI isn't a one time effort**. These systems need to be regularly evaluated by the team because LLM performance can degrade over time. Models change, something's being throttled, context window limits change, and sometimes the cost to performance ratio no longer makes sense. **Maintenance has to be a large part of the planning process, not an afterthought. **If you don't think about maintaining an AI workflow from the moment you first decide that you're going to implement it, you'll end up with either something that nobody uses or something that degrades so much it erodes trust in the system. Going back to our AI code reviewer example, most people set one up within probably a few days through GitHub Actions, a couple of API keys, a couple of calls, it's good, it does what you need. You’re not done, though. Where you're going to win or lose is how you maintain that system and how you make sure that that model is giving consistent results, otherwise all you're doing is creating noise that people won't use. ### Evaluation loops are how you keep the system honest An essential piece of integrated AI workflows is having evaluation loops where you can make sure that the LLM is still performing as expected in each one of these processes. **Implementing these evaluation loops is one of the biggest pieces in productionizing AI.** In practice, evals look like different things depending on the workflow you’re attempting to automate. For an AI ticket intake chatbot, this could be having a generic ticketing workflow that you can automate to pass through the LLM, say once a week. For an automated error handling system, you’d send out a generic error message through your error workflow at a preset cadence so that it can be reviewed and you can make sure that the performance of the AI system you're using stays on point. If you try to skip this part of the process, what can happen is people don't notice the degrading performance at first, until suddenly you're getting skewed data. Again, faulty data is really going to poison any kind of adoption. ### Evals are OG ML Evals aren’t something we’ve just added due to agents. People who've been in the traditional AI frameworks space largely understand that evals are an integral part of the machine learning flow. Before I came to dbt I worked on a very large project doing predictive modeling for a manufacturing company; this was long before agents, but we still did evaluations as part of those workflows. One thing we would do for example is generate a month-over-month capture of basically,_ ‘Hey, we were doing good at predicting X last year. We're not doing so good now. Do we need to retrain or use a different model?’_ A lot of people are coming into AI now because it's the new thing, and often they skip looking into things like evaluating and making sure it's a system that can stand the test of time in production. But it’s always been the case that, whenever you're dealing with a machine doing things on its own, you have to give it guardrails and give it controls, to make sure that you know what's going on under the hood. And I feel like that's missed by a lot of people in the AI race nowadays because everyone's just scrambling to even keep up with what's going on. ### Evals run on humans, too Ultimately, evals are only partly automatable. You could have automated systems that compare the output of your evaluation month over month, but you need a human in the loop for review. If there is any sign of degradation in performance you would want to have a human reviewer apply judgement: is this true degradation, or is the model just giving slightly different answers? Other areas where humans still need to be involved in AI workflows include making decisions around whether to switch models or change something about the skills that you're using in your markdown files or changing what tools the agent can access through MCP. Changes like these need a human to review the results you’re currently getting, and make those determinations around any changes needed, and then rerun those evaluations to make sure that it's still giving the outputs that they would expect My baseline message to builders here is, don't get lazy with it. Don't just accept whatever the AI says, or even use a majority of what the AI says unless it's very boilerplate. Make sure that these outputs are being reviewed, they're being vetted and holding up to the same standards that you've already set for yourself or your team. The advent of AI doesn't mean that we can now just relax on the work that we've been doing for years. ## AI-ready data starts with leadership Right now, AI is very fractured at companies, with different levels of adoption in different areas. Every company wants to adopt AI, and I think most have moved a little too quickly. They give everyone some kind of AI tool and tell them to use it, but what's missing is actually working AI into their _way_ of working. You think you've got AI in engineering because everybody's using Claude Code or Codex, but that's just the top level. That's like saying everybody uses Google at this point. **Making your organization’s data AI-ready means building AI into your workflows.** Not just having your data engineers use AI tools, but also integrating them into your triage flow and your pull request flow and everything else your team does every day. The more a data leader can work AI into the daily workloads of an organization, the more that will actually drive adoption in a meaningful way. This takes buy-in from leadership to say, “When someone submits a pull request, we want to have an automatic agent review. When there's a pipeline failure, we want to be able to see an automatic triage done by the AI system.” For those of you in charge of actually building this, especially when you get those top-down “adopt AI now!” edicts from the C-level, VP level, people who may not be working with this technically — your challenge is communicating what it’s really going to be like. ### The things your org’s leaders need to understand about AI When it comes to fully integrating AI into your organization’s workflows, there are two important pieces of information that need to be communicated up the chain: First, this is extremely exciting technology. But, because AI is non-deterministic by nature, **making AI practical in the enterprise space means making it more deterministic. **You achieve this through evals, and with aligning skills and giving clear instructions and guardrails to make it incrementally more deterministic. As a result, leadership needs to understand that, "Hey, we can't implement this tomorrow." It's like any other major software implementation; actually, it’s even more complex because you're dealing with non-deterministic LLMs. Second, setting realistic expectations is key. Leadership needs to understand that **velocity is not going to 10x overnight.** Velocity will happen, but first the outputs need to be reviewed, the system needs to be trusted, evaluations need to be solid. Developers on the ground floor that are using these AI tools and seeing these outputs for things like error triage, code review, tickets coming in, and even text to SQL, will still need to review the outputs of the LLM. Realistically, you can get efficiency gains with AI, but in these early days, I would say a lot of the efficiency gains you'll get with AI are going to be at least partially eaten up by the amount of work you need to put into making AI reliable and trustable. ## Make AI organizational, not just individual Using AI tools personally is great. It's fantastic, it's super cool, you're spending all your Claude credits, that's awesome. But **if you really want to get the most out of your AI spend and the most out of your AI tooling, you need to work on productionizing it and building it into your daily tasks over time.** That's the muscle we need to be building if we want AI to‌ make sense in our organizations going forward, rather than just as a personal tool. --- --- title: "Data platforms were built to store. Intelligence platforms are built to reason." description: "Data platforms store information. Intelligence platforms govern meaning so AI can reason on it reliably." url: "https://www.getdbt.com/blog/data-platforms-were-built-to-store-intelligence-platforms-are-built-to-reason" date: "2026-07-02" authors: ["Dustin Dorsey"] categories: ["Insights"] --- # Data platforms were built to store. Intelligence platforms are built to reason. _This is a guest post from Dustin Dorsey, Senior Director of Data Engineering at [phData](https://www.phdata.io/) and co-author of Unlocking dbt. He works with enterprise teams to build the data foundations that make AI reliable at scale._ For a long time, the data platform was the destination. You invested in the infrastructure. You built the pipelines. You centralized the data, cleaned it up, made it available. You stood up a warehouse, connected a BI tool, trained your analysts, and delivered dashboards. And for a season, that felt like winning. Because compared to the chaos of disconnected data marts and spreadsheet-driven reporting, it genuinely was. But here is the thing nobody says out loud at the end of that journey: the platform was never the point. The point was better decisions. Faster decisions. Decisions made by people who actually understood what was happening in the business rather than guessing at it. The data platform was supposed to be the foundation for that. In most organizations, it became the ceiling. ## The problem was never infrastructure The data platform era solved the infrastructure problem. Data is more accessible, performant, and widely available than it has ever been. Cloud warehouses eliminated the constraints that used to make analytics painful. Modern ingestion tools reduced the friction of getting data in. Transformation frameworks made it easier to shape data into forms that could be queried and understood. By any technical measure, the infrastructure problem is mostly solved. What was never fully solved is the reasoning problem. The gap between having data and making good decisions from it was always bridged by people. Experienced analysts who understood the business. Data scientists who knew which model outputs to trust and which to question. Executives who had the pattern recognition to know when a number felt wrong. The platform gave these people better tools and faster access. It did not replace the judgment they were applying. For years, that was fine. The platform supported the humans who were doing the reasoning. That was enough because the humans were the bottleneck, and the platform reduced that bottleneck meaningfully. AI changes the equation not because it eliminates human judgment (it does not, and the organizations that believe otherwise are going to discover that the hard way) but because it now sits in the layer where that judgment was always applied. And the platform underneath was never built to support it. An AI system reasoning over your data is not a smarter dashboard. It is a different kind of consumer entirely. Dashboards were built to show humans pre-defined answers that humans then interpreted. AI is expected to form interpretations autonomously. To do that reliably, the platform needs to communicate business meaning, not just store data. And most platforms were never designed to do that. ## The category shift that is actually about philosophy There is a temptation to frame the evolution from data platform to intelligence platform as a technology upgrade. New tools, new capabilities, a few additional layers in the stack. That framing is wrong, and it leads organizations to make the wrong investments. The shift is not about technology. It is about purpose. A data platform is organized around the question: how do we make data available? Every architectural decision flows from that question. Where does data land? How fast can we query it? How do we manage access? How do we ensure it is clean and current? These are the right questions for a platform built to store and serve. An intelligence platform is organized around a different question: how do we make meaning available? The architectural decisions look different when you start from there. It is not enough for the data to be technically accessible. It needs to be interpretable by systems that have no institutional knowledge, no context, no ability to pick up the phone and ask someone what this number actually means. Making meaning available requires decisions that most data platforms deferred. What is the authoritative definition of this metric? What business process does this fact table represent? What does one row in this dataset represent? What relationships between entities are semantically intentional, not just technically possible? These are not new questions. They are questions that were always the right ones to ask. They just did not feel urgent when humans were mediating between the data and the decisions. They are urgent now. ## Overinvesting in intelligence, underinvesting in knowledge Look at where most organizations are putting their AI investment. A large share of it goes into the intelligence layer: the models, the inference infrastructure, the fine-tuning, the orchestration frameworks, the agent architectures. This is where the visible innovation is happening, and it is genuinely impressive. The layer that is chronically underinvested is the knowledge layer. The place where meaning lives. The semantic models, the metric definitions, the ontologies, the shared vocabulary that tells the intelligence layer how to interpret what the data layer contains. This imbalance is understandable. The intelligence layer has vendors, conferences, press releases, and benchmark comparisons. The knowledge layer has unglamorous conversations about what "customer" means and whether that definition should include trial accounts. Organizations gravitate toward work that feels like progress. Getting a model to generate a coherent summary is immediately satisfying. Spending three weeks aligning two teams on a metric definition is not. But the intelligence layer is only as good as the foundation it reasons over. An AI agent that can navigate complex workflows and synthesize information across domains will still produce unreliable outputs if the data it is reasoning over contains five competing definitions of the same concept. The capability of the reasoning layer is bounded by the quality of the knowledge layer beneath it. This is the central mistake of the current wave of AI investment in enterprise organizations: treating the intelligence layer as the place where reliability is established, rather than recognizing that reliability is established in the layers below it and the intelligence layer simply expresses it. ## Meaning has to become a product The organizational shift required here goes beyond architecture. It requires treating enterprise semantics the way mature organizations treat data products. With ownership, investment, stewardship, and a recognition that the value compounds over time. Most organizations do not currently have anyone who owns the definition of their core business metrics. They have people who maintain the pipelines that produce them, people who use them in dashboards, and people who argue about them in meetings. Ownership of the definition, the authoritative, governed, consistently enforced representation of what a metric means and how it should be calculated, is typically nobody's job. For an intelligence platform to function reliably, that has to change. Meaning needs an owner in the same way that a data product needs an owner. Someone who is accountable for whether the definition is current, whether it is consistently applied, whether changes to it are evaluated for downstream impact before they are made. Someone who treats the semantic layer as infrastructure worth maintaining rather than documentation worth archiving. This is not just an engineering problem. It is a product and governance problem. The organizations that recognize this early will build the organizational muscle to maintain their knowledge layer the same way they maintain their data infrastructure. The ones that treat it as a one-time modeling project will find themselves rebuilding it repeatedly as the business changes and the accumulated semantic drift undoes the previous round of work. ## What an intelligence platform actually looks like Calling something an intelligence platform is easy. The organizations doing it seriously are making a specific philosophical commitment, and it is worth naming clearly. A data platform asks: what do we have and where does it live? An intelligence platform asks: what does it mean and how should it be interpreted? That sounds like a semantic distinction until you see how differently the two platforms behave when an AI system is reasoning over them. The data platform produces a technically accessible answer. The intelligence platform produces the right one. The difference lives in a layer that most data platforms treated as an afterthought: the place where business meaning is governed. Not catalogued since catalogues tell you where data lives and what columns it contains. Governed. Where the definition of a metric is owned, maintained, and enforced in the transformation layer rather than documented in a wiki and applied inconsistently. Where the question "what does one row in this table represent" has an answer baked into the structure itself rather than living in an analyst's institutional memory. Organizations that try to build AI capabilities without that layer will keep encountering the same problem in different forms. They will tune the model and get better outputs for a week. They will refine the prompts and get consistency on the queries they thought to write prompts for. But the underlying ambiguity will surface somewhere else, because you cannot reason reliably over a foundation that was never designed to be reasoned over. The layer that makes AI reliable is not the AI. It is what the AI is built on top of. ## In five years, "data platform" will sound like a storage description Predictions about technology timelines are usually wrong in the specifics and right in the direction, so take this with appropriate skepticism. But the direction feels clear. The organizations building intelligence platforms today are not building something entirely new. They are building what data platforms were always supposed to be. The original promise of the data platform was that better data access would lead to better decisions. That promise was partially fulfilled. The AI era is forcing the other half to be taken seriously. The half about making meaning accessible, not just data. As that shift happens, "data platform" will increasingly feel like a description of what something stores rather than what it does. The platforms that persist and grow will be the ones organized around meaning and reasoning, not just storage and query. The investments that compound will be the ones made in the knowledge layer. In canonical definitions, in intentional data models, in the semantic infrastructure that makes it possible for both humans and AI systems to reason over enterprise data consistently. The organizations that are already making those investments do not look dramatically different from the outside today. Their AI demos are not necessarily more impressive than anyone else's. But when they move those demos into production, the outputs hold. The answers are consistent. The metrics mean the same thing to the finance team and the sales team and the AI system querying the data at 2 in the morning. That consistency is not a technical artifact. It is the result of someone having made deliberate choices about meaning and encoded them into the platform itself. That is what an intelligence platform actually is. Not a new set of tools layered on top of an existing data stack, but a fundamentally different philosophy about what the stack is for. And a recognition that the data engineering discipline has always been building toward this, even when the immediate deliverable was just a faster pipeline. If this series has prompted you to rethink what your own data foundation needs to look like for AI to operate reliably on top of it, [Building the Foundational Layer for Reliable AI on Structured Data ](https://www.phdata.io/offers/ai-data-foundation-whitepaper/)is the place to go deeper. It covers the structural conditions that separate AI environments that scale reliably from the ones that keep surfacing the same unresolved problems and why dimensional modeling is the prerequisite that most organizations are skipping. ## Where dbt and phData fits dbt sits at exactly the inflection point this shift requires. As organizations move from data platforms to intelligence platforms, the transformation layer becomes the place where meaning gets encoded and enforced, and dbt provides the model structure, semantic governance, and integration surface to make that possible. phData helps teams operationalize that shift by translating the philosophy into an executable data foundation: designing the knowledge layer, implementing the supporting models and semantics, and helping organizations move from ideas about reasoning to systems that support it in practice. --- --- title: "What Fivetran + dbt Labs brought to Databricks Data + AI Summit (and what you can take home)" description: "See what Fivetran + dbt Labs shared at Databricks Data + AI Summit: dbt Wizard, dbt State, and dbt Core v2.0, plus what's next." url: "https://www.getdbt.com/blog/what-fivetran-dbt-labs-brought-to-databricks-data-ai-summit-and-what-you-can-take-home" date: "2026-06-23" authors: ["Daniel Poppy"] categories: ["Partnerships"] --- # What Fivetran + dbt Labs brought to Databricks Data + AI Summit (and what you can take home) If agents are becoming a primary consumer of your data, what does that mean for your data work? It’s a lot to think about, and a lot to talk about, and ask questions about, which is exactly what we heard at the 2026 Databricks Data + AI Summit. Agents are only as reliable as the governed, tested, traceable data beneath them, which is why we showed up in San Francisco last week to show what’s important to us: data infrastructure for agents you trust. [We recently announced that dbt Labs and Fivetran are one company](https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means). It makes Databricks Data + AI Summit all the more meaningful to us to show up with our shared mission to build Open Data Infrastructure for the agentic era. ## What we shared at Databricks Data + AI Summit ![photograph of dbt staff at Databricks Data + AI Summit](https://cdn.sanity.io/images/wl0ndo6t/main/00b2de5e789d2a58e99a256743abea59a46e45c2-4284x5712.jpg) We had a great time at Databricks Data + AI Summit, and if you were there, we hope you had a chance to chat with us. Based on the number of folks that attended our sessions, visited our booth, or attended our events,‌ a lot of you tried. And if you didn’t have a chance, we’d love to talk with you about how you’re approaching the data and agents landscape. [Talk to a dbt expert now](https://www.getdbt.com/contact). Kicking things off, Fivetran + dbt Labs CEO George Fraser and Fivetran + dbt Labs President Tristan Handy took the stage to share their vision of how AI is reshaping data infrastructure for the next decade. As agents become a primary consumer of data, they query the warehouse continuously, which raises the stakes on cost, observability, and above all, context. For George and Tristan, the answer to this is Open Data Infrastructure. The context your agents need lives in your data platform. For data teams, that means providing trusted, governed context for AI becomes a core part of the job. [To get direct access to Tristan and Fivetran + dbt Labs COO Taylor Brown, join our upcoming webinar and ask any questions you have about the new company.](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a) ## dbt and Fivetran customers sharing their journey ### The “BI why” problem Our customers provide an on-the-ground experience of what it’s like for data teams right now. Mars Dauer, Senior Director of Enterprise Data and AI at Integral Ad Science, spoke about what he calls the “BI why” problem: when a BI copilot can’t tell you why a number is wrong. It happens because the copilot only sees the top of the stack, and not the business logic underneath that produced the number. His team’s solution is an AI agent that connects Databricks and dbt through the model context protocol (MCP). The agent draws on three MCP servers, the dbt MCP server for model metadata, lineage, and compiled SQL, a Databricks SQL MCP server for read-only validation governed by Unity Catalog, and a Looker MCP server as the BI on-ramp. What used to take an analyst an afternoon, now takes minutes, Dauer says, and they are only scratching the surface. We recently announced [dbt Wizard, an agent built for analytics engineers](https://www.getdbt.com/product/dbt-wizard). It works alongside you with deep domain knowledge and a complete understanding of your project and workflows. [Join us for a deep dive to demo dbt Wizard in two surfaces: the CLI for engineers who live in the terminal, and the dbt platform for teams that want a shared interface](https://www.getdbt.com/resources/webinars/dbt-wizard-an-agent-purpose-built-for-analytics-engineering). ### Scale data and AI with Fivetran, dbt, and Databricks Inova Health Chief Data + AI Officer Jon McManus shared how the healthcare provider compressed a 4-year modernization roadmap into 6 months with Fivetran, dbt, and Databricks. [His team cut data movement costs by up to 8x with Fivetran](https://www.fivetran.com/case-studies/inova-health-compresses-4-year-roadmap-into-6-months-to-power-ai). dbt introduced a shared layer for transformation and [governance](https://www.fivetran.com/governance), which enables the team to define consistent metrics and improve trust in data across the enterprise. And Databricks provided a scalable platform to support advanced analytics and AI workloads. ## dbt on display at Databricks Data + AI Summit We’ve announced a lot recently, and it was pure joy giving folks their first look at everything in our booth and sessions and hands-on labs. With Fivetran and dbt Labs now as one company, we’re building the open data foundation for the agentic era so that your business logic and governance travel with you. We loved talking with folks new to dbt as they discovered the power of the dbt platform. And for current dbt users, we got to show off what’s new: [**dbt Wizard**](https://www.getdbt.com/product/dbt-wizard) is an AI agent purpose-built for analytics engineering to understand your lineage, tests, contracts, and metric definitions. [**dbt State**](https://www.getdbt.com/product/dbt-state) skips unchanged models across production and local runs for 30% average warehouse compute savings. And we announced the [**first alpha release of dbt Core v2.0**](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion), now built on the same foundations as the dbt Fusion engine. dbt Core v2.0’s key feature developments include: - **Significant parse time improvements** - **A tightly-defined language spec** that gives anyone integrating with the dbt ecosystem a stable interface to build against. - **New Parquet artifacts** as a high-performance alternative to large JSON files that can be directly queried through your [agent of choice](https://www.getdbt.com/product/dbt-wizard). - **A completely revamped local documentation experience,** powered by those new artifacts and capable of scaling. - **A more streamlined way to build new adapters**, powered by Arrow Database Connectivity (ADBC) and the Arrow ecosystem. - **A simplified installation process** that removes the need to fight with Python's virtual environments. ## The dbt community brings the energy Whether you stopped by our booth, attended our talks, or joined our Happy Hour by the Bay or executive dinners, we’re grateful to connect with you. The dbt community is doing the most exciting work in data. If you want to see how the dbt community REALLY shows up, [**join us at dbt Summit Sept. 15-18 in Las Vegas**](https://www.getdbt.com/dbt-summit). dbt Summit is the world’s largest gathering of dbt users. Level up your data and AI work with dbt community members to shape the future of analytics engineering. Whether it’s at [dbt Summit](https://www.getdbt.com/dbt-summit), our [webinars](https://www.getdbt.com/resources/webinars), a [dbt meetup](https://www.getdbt.com/events), or [dbt Slack](https://www.getdbt.com/community/join-the-community), we’ll see you soon. ```json { "_key": "acdecba82c76", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/PcLd5bDRanQ" } ``` --- --- title: "The semantic debt crisis no one is talking about" description: "Two teams. Same metric. Different numbers. That's semantic debt, and AI will make it impossible to ignore." url: "https://www.getdbt.com/blog/the-semantic-debt-crisis-no-one-is-talking-about" date: "2026-06-22" authors: ["Dustin Dorsey"] categories: ["Insights"] --- # The semantic debt crisis no one is talking about _This is a guest post from Dustin Dorsey, Senior Director of Data Engineering at phData and co-author of Unlocking dbt. He works with enterprise teams to build the data foundations that make AI reliable at scale._ Picture the scene. Two teams are presenting to the same leadership group. Same company. Same time period. Same metric. Different numbers. Both teams defend their number with confidence. Both can show their work. Both are, by their own definition of the metric, correct. The meeting derails. Someone from finance says the sales team's number is wrong. The sales team says finance is using the wrong date field. Someone points out that marketing has a third version that does not match either of them. An hour passes. Nothing gets decided. The executive sponsor asks someone to "get aligned on the numbers" before the next meeting, and everyone leaves knowing that conversation will take months if it happens at all. This is not a data quality problem. The data is accurate. The pipelines ran correctly. The numbers are precisely what they claim to be. It is a design problem. Business meaning was never encoded into the data itself, so every team built their own version of the truth, and now there are three of them. Each defensible, none authoritative, all quietly costing the organization more than anyone has ever calculated. __ _Dustin goes deeper on the data design decisions that make AI reliable in production: [Your AI isn't broken. Your data model is.](https://www.getdbt.com/blog/your-ai-isn-t-broken-your-data-model-is)_ ## This is what semantic debt looks like Technical debt is a concept most engineers know well. You take a shortcut to ship faster, and you pay for it later in maintenance, refactoring, and fragility. Semantic debt works the same way, except the thing being deferred is not code quality. It is shared meaning. Every time a team builds a metric without encoding a canonical definition, that is a withdrawal from the semantic account. Every time a dashboard hardcodes a filter that belongs in the data layer, that is a withdrawal. Every time a business concept gets defined slightly differently in two separate pipelines, both of which work fine in isolation, that is a withdrawal. No individual shortcut looks catastrophic. The accumulation is. In most organizations, semantic debt has been building for years, sometimes decades. It was built quietly, because humans are remarkably good at compensating for it. A finance analyst knows their "revenue" is not the same as the sales team's "revenue" and calibrates accordingly. A data engineer knows which version of a customer record is authoritative for which use case. A business intelligence developer knows to add that one filter that nobody can explain but everyone agrees must be there. These workarounds are so ingrained they stop feeling like workarounds. They feel like just how things work. What they actually are is institutional knowledge plugging structural gaps that should have been designed out years ago. ## The meeting AI cannot call In a human-driven analytics environment, semantic debt gets resolved through a specific, recognizable process. Two teams disagree on a number. Someone calls a meeting. In the meeting, each team explains their logic. The group agrees (usually informally) on which definition is correct for which purpose. Someone writes a note in a Confluence page that nobody will find in two years. The organization moves on, with a slightly better shared understanding that lives, once again, in people rather than systems. This process is inefficient. It is slow. It only surfaces the conflicts visible enough to generate a meeting. But it more or less works, because the scale of human-driven analytics is limited enough that the gaps can be managed. AI cannot call that meeting. When an AI system is asked a question that depends on a metric your organization has defined inconsistently, it does not pause, identify the conflict, schedule a working session, and wait for consensus. It picks an interpretation and produces an answer. Confidently. Without flagging which version of the definition it used. Without knowing that a different version exists. Without any signal to the downstream consumer that the answer may depend on a definitional assumption that three teams would argue about if they knew it was being made. The conflict that your organization has been managing through meetings and tribal knowledge for years does not disappear when AI arrives. It gets made at machine speed, at scale, without anyone in the room to catch it. ## Where meaning lives in most organizations If you were to trace the authoritative definition of your organization's most important metrics, you would find them in some combination of the following places: a Confluence page last updated eighteen months ago, a comment in a SQL file that three engineers have each interpreted slightly differently, the institutional memory of a senior analyst who has been with the company for eight years, a Slack thread from 2022 that was pinned for a while and then wasn't, and a spreadsheet that lives on someone's desktop that nobody else knows about. This is not an exaggeration, and it is not a sign of dysfunction. It is the default state of organizations that built their data stacks to move fast and deliver value quickly within specific domains. Meaning accumulated in people because people were always there to apply it. Encoding it into the data structure itself was extra work that did not obviously improve the metric that mattered, which was whether the dashboard loaded. The problem with meaning living in people is not that people are unreliable. It is that meaning needs to travel to places people cannot go. It needs to travel to the AI system that a business user is querying at 11pm trying to understand why their numbers look different from last week. It needs to travel to the automated pipeline that is making decisions based on customer behavior without a human in the loop. It needs to travel to the new analyst who joined the company six months ago and has never heard the story of why that filter is always set to exclude refunds. It needs to travel to every downstream consumer of the data, reliably and consistently, every single time. People-stored meaning cannot do that. Only data-encoded meaning can. ## The stakes are higher than they appear Most organizations experience semantic debt as an inconvenience. Meetings that take longer than they should. Reports that require manual reconciliation before they can be shared. Analysts who spend a third of their time validating numbers before they trust them enough to use. These costs are real, but they are diffuse and largely invisible on any financial statement. AI changes the cost structure entirely. When AI is operating on data where meaning is inconsistently defined, the outputs are not just inconvenient, they are unreliable at a scale no human team could produce. An analyst who does not know which revenue definition to use asks a clarifying question. An AI that does not know which revenue definition to use picks one and answers three hundred queries with it before anyone notices the problem. The other shift is who is affected. Traditional analytics tools were used primarily by trained practitioners who understood the limitations of the data they worked with and knew, from experience, where the landmines were. AI-powered interfaces lower the barrier dramatically. Business users who have never seen the underlying data model, executives who expect answers to just be right, external stakeholders in some cases, all of them are now interacting with data whose meaning was never designed to be self-evident. Semantic debt that was manageable when experts were the only users becomes genuinely dangerous when anyone with a natural language prompt can query your data directly. ## You have a preview, not a problem Here is the argument I want to make directly: if your most important business metrics mean different things to different teams today, you do not have an AI problem yet. You have a preview of one. The meeting where two teams present different numbers? That is semantic debt becoming visible. It is uncomfortable and inefficient, but it is still being caught. Someone is in the room. The conflict surfaces. It gets resolved, however imperfectly. The question is whether your organization addresses the underlying cause intentionally (before AI amplifies it) or reactively, after AI has already made it impossible to ignore. Reactive is expensive. Not just in technical remediation time, but in organizational trust. The erosion of confidence in AI outputs is one of the hardest things to reverse once it begins. Business users who receive inconsistent answers stop trusting the system. Executives who got burned by a bad AI-generated number become skeptical of every AI-generated number. The technology gets blamed for a problem that the technology did not create, and the actual cause (the accumulated semantic debt in the data layer) stays unaddressed because the room full of people who saw the symptom did not get far enough upstream to see the disease. The organizations that act now have a significant advantage. Not because they will finish building perfect semantic foundations before anyone else because that is not how this works, and the organizations claiming they will are deceiving themselves. The advantage is that they will start building them as a deliberate strategic investment rather than as an emergency response to a crisis that has already damaged trust. ## Making meaning a system property, not a people property Think about what it would actually take for a new analyst to answer the question "what is our monthly active user count" correctly in their first week. They would need to know which table is authoritative, which definition of "active" the business uses, which date field governs the period, which user types to exclude, and probably the history of why each of those decisions was made. None of that is in the data. It lives in people. The new analyst calls a meeting. Someone explains it. They write it down in a place only they will ever look. Two years later, a different analyst asks the same question and starts over. This is the loop that data-encoded meaning breaks. When business logic is enforced in the structure itself (in how grain is declared, in how relationships are modeled, in what the transformation layer enforces) a new analyst and an AI system get the same correct answer without needing to know the history. Not because someone documented it well. Because the structure made the wrong answer impossible to produce. The distinction matters because meaning needs to travel. It needs to travel to the new analyst who joined six months ago, to the automated pipeline running without human supervision, to the AI system answering questions at 11pm. None of those consumers can call a meeting when they encounter ambiguity. They either get a consistent answer because the structure provides one, or they get a defensible guess because it does not. This is also not a one-time project. It is ongoing stewardship. Meaning drifts as the business changes. New products launch. Definitions evolve. Mergers happen. The organizations that stay ahead of semantic debt are not the ones that did a modeling exercise in 2022 and called it done. They are the ones that treat their semantic foundation the same way they treat their data pipeline. As infrastructure that requires ownership, investment, and deliberate evolution. What most organizations are missing is not the technical ability to do this. It is the organizational will to treat it as infrastructure rather than overhead. That shift (from viewing semantic work as a documentation task to viewing it as a platform investment) is what separates the organizations that will navigate the AI era well from the ones that will spend the next several years explaining why the answers keep changing. ## What you should be asking right now The diagnostic for semantic debt is not a technical audit. It is a business conversation. Pick your organization's three most important metrics. Ask five people from different teams to define them. Not to query them, but to define them. To explain what counts, what does not, what edge cases apply, and how they should be calculated. If you get the same answer from all five people, you are in better shape than most. If you get five different answers, each defensible, you are looking at your AI risk surface. The follow-up question is where the decision lives: do you want to address this before the AI initiatives your organization is investing in expose it at scale, or after? There is no neutral choice here. Semantic debt compounds. Every month it goes unaddressed is another month of diverging definitions, additional teams building their own versions of the truth, and growing distance between where your data is and where it needs to be for AI to operate reliably on top of it. Addressing semantic debt is not just about fixing individual models, though. It requires rethinking the fundamental purpose of the data platform itself. Most organizations are building platforms designed to store. What the AI era requires is platforms designed to reason. That is the argument in the next post and it is a bigger shift than it sounds. If the patterns in this post feel familiar, [Building the Foundational Layer for Reliable AI on Structured Data](https://www.phdata.io/offers/ai-data-foundation-whitepaper/) goes deeper on the structural conditions that separate data environments where AI operates reliably from the ones where it keeps surfacing the same unresolved ambiguity. Worth reading before your next AI planning conversation. ## Where dbt and phData fit dbt is one of the few tools in the modern data stack where semantic debt can be paid down incrementally and sustainably. Its combination of model-layer definitions, the dbt Semantic Layer for centralized metric governance, and documentation gives teams a practical path to encoding meaning as part of how transformations are written and maintained. phData helps teams put that into motion by working through the harder operational pieces alongside dbt: clarifying definitions, reshaping models, and building the delivery discipline needed to make semantic consistency hold up beyond a single project or team. --- --- title: "Start fresh, don't lift and shift: a dbt migration guide" description: "dbt migrations underdeliver when teams rebuild legacy patterns in a new tool. Here's how to start fresh instead." url: "https://www.getdbt.com/blog/start-fresh-don-t-lift-and-shift-a-dbt-migration-guide" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Start fresh, don't lift and shift: a dbt migration guide We've seen this pattern often enough to name it. A team migrates to dbt, spends six months, and ends up with a dbt project that looks exactly like their old workflow, just with Jinja templating instead of drag-and-drop. The data model still has the same problems. It's only newer. The migration project gets checked off as complete. The downstream problems show up six months later: reports that contradict each other, engineers who can't explain what a model does without reading the SQL, a semantic layer that makes data trust worse instead of better. This is the lift-and-shift problem, and dbt isn't the cause of it. ## Signs your dbt project is a legacy migration The symptoms are recognizable once you know what to look for. Here are the six most common. **Every stored procedure has a 1:1 dbt model equivalent.** The migration team mapped source code to dbt models one-for-one. The logic lives in SQL now instead of PL/SQL, but the structure is identical, and no one asked whether the original structure was worth keeping. **Models are named after source systems, not business entities.** You see `salesforce_accounts_cleaned` and `hubspot_contacts_deduped` instead of `customers` and `leads`. The names describe where data came from rather than what it means to the business. **No staging models exist.** Everything jumps directly from raw source to reporting table. The [dbt project structure guide](https://docs.getdbt.com/best-practices/how-we-structure/1-guide-overview) is explicit: [staging models](https://docs.getdbt.com/best-practices/how-we-structure/2-staging) are the atomic building blocks of a dbt project, and skipping them skips the part that makes the architecture maintainable. **Tests are bolt-ons rather than design principles.** The migration got done, then someone added [`not_null` tests](https://docs.getdbt.com/docs/build/data-tests) to the most important columns as an afterthought. **Documentation is empty or copied from source-system field descriptions.** Column descriptions say things like "from MDM system." That records provenance, which is useful, but it leaves the actual documentation undone. **No `ref()` lineage.** Models depend on hardcoded schema names instead of [`ref()`](https://docs.getdbt.com/reference/dbt-jinja-functions/ref). When something breaks, teams find out from a failed report rather than a CI check. None of these is hard to fix individually. The problem is what they signal together: the team migrated code, and left the thinking behind. ## The rewrite feels risky and the migration feels safe. Look closer. Here's how teams end up here. Migration projects get measured by completion, not quality. There's a deadline, there's a checklist, and the checklist says "migrate 200 stored procedures." Whether those procedures represent good data modeling never makes the checklist. The path of least resistance is to replicate what exists. Translate the stored procedures, get to green on the migration tracker, and ship. Quality gets deferred, and deferred quality hardens into permanent technical debt. A recent consulting engagement shows the pattern. An insurance company undertook a major migration from a legacy ETL platform to dbt. The migration ran on schedule and on budget. A year later, the data team was spending more time debugging model failures than shipping new analytics. The architecture had carried over all the original fragility, in a different tool. The symptoms: 300-plus models with no staging layer, no tests on roughly 40% of models, reports built directly on raw source joins, and a semantic layer that three different teams maintained independently. The migration was complete. The architecture was broken. Here's the framing that helps. A lift-and-shift migration is really a replatform: the tool changed, and the system stayed the same. If the data model was wrong before dbt, it's still wrong after. ## The triage decision: which legacy models deserve to exist Before writing a line of dbt SQL, teams should ask a harder question than "how do we migrate this?" The better question: "Does this deserve to exist?" Here's a practical triage framework. **Eliminate.** Reporting tables built for a deprecated BI tool, summary tables that exist only because the old database couldn't handle the underlying query, and "just in case" tables nobody queries. Leave these behind. **Rewrite.** Any logic that joins raw source tables with no staging layer. Any model that does transformations and aggregations in the same step. Any model named after source systems rather than business entities. Redesign these rather than translating them. **Translate with care.** Validated, tested business logic the organization depends on. Bring it over with explicit [tests](https://docs.getdbt.com/docs/build/data-tests), document every column, and have someone who understands the business domain review the logic, not just the SQL. **Build fresh.** The [semantic layer](https://www.getdbt.com/product/semantic-layer). This should always be built from business requirements: what is a customer, how do we define revenue? Migrating legacy SQL into the semantic layer inherits every ambiguity and compromise baked into the original definitions. The migration health conversation with stakeholders usually centers on timeline and scope. The more useful conversation is about triage: what are we keeping, what are we rebuilding, and what are we cutting? ## Seven signs of a healthy dbt migration Use this as a diagnostic. If you're mid-migration, run through it this week. 1. **Every source model has a [staging model](https://docs.getdbt.com/best-practices/how-we-structure/2-staging).** The staging model cleans and standardizes the data: type casting, column renaming, basic validation. Nothing downstream touches raw source tables directly. 2. **Business entities are represented as marts.** Final models are named for business concepts: a `customers` mart, an `orders` mart, a `revenue` mart. They represent what the business cares about, not what the source systems happen to contain. 3. **Every model has at least [`not_null` and `unique` tests](https://docs.getdbt.com/docs/build/data-tests) on primary keys.** These are the minimum, and they catch the most common failure modes: duplicate rows and unexpected nulls. Without them, you have data hope, not data quality. 4. **Documentation coverage is tracked and improving.** Not every column needs a long description, but every model should have one that tells a new team member what it is and why it exists. 5. **[`ref()`](https://docs.getdbt.com/reference/dbt-jinja-functions/ref) is used everywhere.** No hardcoded schema names. If the project can't run in a fresh schema without breaking, it isn't production-ready. 6. **CI runs on every PR.** Every pull request runs tests before merge, so nobody merges broken SQL. 7. **A team member who didn't write a model can explain what it does from the documentation.** If someone can't understand a model without reading the underlying SQL, the documentation has failed, and eventually the model will too. ## Why migration quality matters even more in 2026 The pressure to migrate from legacy systems to dbt is real, and so is the pressure to do it fast. In 2026, a third pressure has arrived: the AI systems being built on top of data infrastructure are only as good as that infrastructure. A lift-and-shift migration produces exactly the kind of foundation that makes AI unreliable. Models without semantics, tests, or documentation mean AI systems inherit every gap. The semantic layer doesn't know what "customer" means. Tests don't exist to catch when something breaks. Documentation can't help a model trace back to its source. The data shows how wide the gap is. Gartner [predicted in early 2025](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) that through 2026, organizations would abandon 60% of AI projects built on data that isn't AI-ready. And according to a March 2026 report from Cloudera and Harvard Business Review Analytic Services, only 7% of enterprises say their data is completely ready for AI, while 73% say their organization should prioritize AI data quality more than it currently does. The teams building reliable AI systems are doing it on data infrastructure they actually trust. That starts with the migration, well before the AI project. ## Closing Two practical moves from here. If you're mid-migration right now, run the seven-sign checklist as a diagnostic this week. Don't wait until the migration is "done." The checklist will tell you how far the project has drifted from good architecture and how much rework is piling up. If you're planning a migration, the triage framework is the architecture conversation to have before writing a line of code. Get the team in a room, look at the models in scope, and answer one question: does this deserve to exist in the new system? --- --- title: "The analytics engineer in 2026: system designer, governance owner, AI context provider" description: "AI is reshaping the analytics engineer role. Here's what system design, governance, and AI context look like in 2026." url: "https://www.getdbt.com/blog/the-analytics-engineer-in-2026-system-designer-governance-owner-ai-context-provider" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # The analytics engineer in 2026: system designer, governance owner, AI context provider ## What analytics engineering looked like in 2023 In 2023, the core of an analytics engineer's job was model development. You wrote SQL, organized it into dbt models, wrote tests, and built pipelines that turned raw data into something stakeholders could use. Documentation was a best practice you aspired to. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) was a nice-to-have. The bottleneck was your capacity to write and review code. In 2026, that bottleneck has eased. AI can write dbt model scaffolding faster than any human. It can generate first-draft documentation from lineage metadata. It can write the boilerplate tests most models need. AI-assisted coding is now part of how most analytics engineers work: 72% of them, according to the [2026 State of Analytics Engineering report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026). The repetitive parts of model production are increasingly automated. That clarifies the role rather than shrinking it. With the repetitive work no longer the bottleneck, what's left is the work analytics engineers were always most valuable for. ## The three new responsibilities of the analytics engineer in 2026 **System design** The analytics engineer in 2026 focuses less on individual model implementation and more on how the system of models works. Which models are the source of truth for which metrics? Where are the boundaries between domains? How should the [semantic layer](https://www.getdbt.com/product/semantic-layer) be structured so downstream AI queries return consistent answers? These are architecture decisions that require business judgment and an understanding of how the data gets used, not just how it gets built. AI can scaffold a model. It can't decide whether revenue should be defined at the order line level or the order level, or which grain is correct for a retention metric. That judgment requires understanding the business, which remains a human capability. (For a real-world look at the tradeoffs, see [who should own the semantic layer](https://www.getdbt.com/blog/semantic-layer-ownership).) **Governance ownership** As AI-assisted development accelerates data production, the governance layer becomes more important. Tests, [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts), column-level lineage, and ownership assignment are now the outputs that separate a trustworthy data system from a fast but unreliable one. The analytics engineer owns those outputs. This changes how the role gets evaluated. In 2023, an analytics engineer's output was models. In 2026, it's also the [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) that protect those models, the [tests](https://docs.getdbt.com/docs/build/data-tests) that validate them, and the semantic definitions that make them machine-readable. Governance has become a primary deliverable. (More on this in [semantic layer for data governance and security](https://www.getdbt.com/blog/semantic-layer-data-governance-security).) **AI context provision** This responsibility has emerged most visibly in the past eighteen months, and it's the one analytics engineers are often not trained for explicitly. AI agents need context to reason reliably, and [that context has to come from somewhere](https://www.getdbt.com/blog/how-a-semantic-layer-prevents-ai-hallucinations-in-analytics). In a well-structured dbt project, it comes from [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definitions, column-level lineage, model documentation, and schema contracts. The analytics engineer who understands how to structure that context, what to name things, how to document them, which definitions to make machine-readable, is directly improving the reliability of every AI agent that runs against the data stack. The data community has spent a lot of 2026 debating "context" as a buzzword, and the joke is fair. The underlying problem it names is real: organizations are failing at AI not because their models are wrong, but because AI agents lack the context to reason about what the data means. Analytics engineers build that context. That's a significant expansion of the role's leverage. ## Analytics engineer skills that matter in 2026 SQL fluency still matters, and it isn't going away. But the premium on raw SQL productivity is lower than it was two years ago, because AI can produce syntactically correct SQL faster than most humans. Three things have gotten more valuable: business judgment, semantic precision, and system thinking. Business judgment means understanding what the data represents well enough to know when an AI-generated model is wrong, even when it looks syntactically correct. It means knowing that a metric definition that works for one use case will mislead in another. That judgment isn't automatable. Semantic precision means writing metric definitions and documentation precise enough to be unambiguous, both to a human reading them and to a model reasoning over them. This is a new skill, and the analytics engineers who develop it are more valuable in an AI-native data stack. System thinking means understanding how models relate, where the dependencies are, and how architectural decisions propagate through the stack. As AI takes over individual model implementation, the analytics engineer's comparative advantage lies in the decisions that span models. ## The career case for analytics engineers in 2026 Analytics engineers are more valuable in 2026 than they were in 2024, and AI is the reason why. In 2024, some analytics engineering work created value and some was maintenance. AI is eliminating the maintenance, and what remains is disproportionately the value-creating work: semantic definition, governance ownership, architecture decisions, and the business judgment that determines whether fast data is also accurate data. An analytics engineer who spends 2026 competing with AI on code production will find the role shrinking. One who focuses on what AI can't do, business judgment, context design, and governance ownership, will find it expanding. Treat that as a specific description of what to build toward. ## How dbt supports the evolving analytics engineer role dbt is well-positioned for this shift because of what it has always stored: semantic context in code. The metric definitions, tests, contracts, and documentation in a dbt project are exactly the context AI agents need to reason about data reliably. The analytics engineer who maintains that context is the person making AI-assisted data work trustworthy. The tooling supports this directly. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) gives analytics engineers an AI agent grounded in their project's lineage, contracts, tests, and metrics. The [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) makes the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) queryable in natural language. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) gives AI agents the provenance they need to cite their answers. None of that replaces the analytics engineer. It amplifies what the analytics engineer already does: define what data means, govern how it's used, and build the context AI depends on. That work is the most valuable in the data stack right now, and the question is whether analytics engineers recognize it as such. --- --- title: "How dbt makes agentic data pipelines trustworthy: the transformation layer's role in autonomous data systems" description: "AI agents can run your pipelines. But who decides what \"correct\" looks like? The transformation layer does, and that layer is dbt." url: "https://www.getdbt.com/blog/how-dbt-makes-agentic-data-pipelines-trustworthy-the-transformation-layer-s-role-in-autonomous" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # How dbt makes agentic data pipelines trustworthy: the transformation layer's role in autonomous data systems ## What's missing from every agentic data pipeline diagram Agentic, self-healing pipelines are having a moment. Dagster has published its [AI-driven data engineering vision](https://dagster.io/blog/announcing-ai-driven-data-engineering). The Airflow community is discussing [agentic workloads in Airflow 3](https://airflow.apache.org/blog/agentic-workloads-airflow-3/). Datafold's [2026 predictions](https://www.datafold.com/blog/data-engineering-in-2026-predictions/) put autonomous data engineering on the near-term roadmap. The architecture diagrams all show the same thing: sources feeding agents that run tasks that produce results. None of them shows the layer that determines whether those results are correct. That layer is the transformation layer. And the question nobody seems to be asking yet: when an AI agent builds and runs a data pipeline autonomously, who defines what correct looks like? Without an answer, a self-healing pipeline is just a fast pipeline. It heals quickly, and it propagates wrong answers quickly. Speed isn't the value here. Correctness is, and correctness requires a governed transformation layer. ## What the transformation layer does in an autonomous system In a human-operated pipeline, the transformation layer is where raw data gets shaped into something meaningful. Models define how tables relate. Tests assert that specific conditions hold. Contracts enforce that the shape of data at a boundary can't change without an explicit decision. In an agent-operated pipeline, all of that still has to happen. The difference is that the agent making the changes doesn't inherently know your business rules. It knows syntax. It knows patterns from training data. It doesn't know that your revenue metric must exclude refunds, or that a customer is only active if they've logged in within 30 days, or that a null in this column means something different than a null in that one. That knowledge has to be encoded somewhere. In dbt, it lives in models, tests, [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts), and [semantic layer](https://www.getdbt.com/product/semantic-layer) metric definitions. Encoding it is the core of what dbt does. (For more on this, see [what agentic AI requires from your data](https://www.getdbt.com/blog/agentic-ai-data-requirements).) When an AI agent operates a pipeline with a governed dbt transformation layer, it isn't making autonomous decisions about what the data means. It executes transformations whose semantics were defined by humans, validated by tests, and protected by contracts. The agent gets speed. The business gets correctness. That's the value of governed agentic workflows. ## Why governance has to come before autonomy Teams that skip the governance step and connect AI agents directly to their transformation layer are automating the wrong thing. An AI agent running against a dbt project with no [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) can change a column type in a core model and break every downstream metric silently. An agent running against a project with no [semantic layer](https://www.getdbt.com/product/semantic-layer) definitions interprets metric names however seems reasonable from the table structure. Sometimes it's right. Often it's confidently wrong in ways that are hard to debug. The pipeline self-heals. The numbers are still wrong. The CFO doesn't care that the pipeline ran without errors. Governance at the transformation layer is the prerequisite for AI autonomy, not a constraint on it. An agent that can trust the semantic definitions it works with operates faster and with more autonomy, because the boundaries are what make autonomy safe. ## How data contracts and live project context create a trusted layer dbt gives agentic workflows two things that matter most: trusted boundaries and current context. [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) define the shape of data at a model boundary: which columns must exist, what types they must be, which constraints must hold. When a contract is defined, breaking changes are caught at compile time, before they run. An AI agent that tries to remove a column a downstream contract depends on produces a compilation error, not a silent pipeline failure. Context matters just as much. An agent working from stale manifests reasons about a project that may be hours out of date. dbt's metadata layer and [column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) in dbt Explorer give agents accurate column types, current dependencies, and real lineage, so an agent writing code is working from the actual state of the project. Together, these give an agent a context it can trust. It knows what the columns mean. It knows where the boundaries are. It knows that crossing one without an explicit decision means the pipeline won't run. That structure is what makes real autonomy possible, rather than uncontrolled automation. ## MetricFlow: defining what "correct" means for AI agents Contracts protect structure. The semantic layer defines meaning. [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definitions encode the business logic that turns data into answers. Revenue isn't just a `SUM(amount)` column. It's a `SUM(amount)` with specific filters, from specific models, under specific conditions, for specific purposes. That definition, written once in MetricFlow and version-controlled in the project, is the canonical answer to "what is revenue?" When an AI agent queries the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) through the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), it isn't making up an answer. It queries a definition a human wrote, reviewed, and committed. The answer holds whether the question comes from a BI tool, a Slack bot, or an autonomous agent running a scheduled pipeline. That's what "correct" looks like in an autonomous data system: a model that returns the answer your CFO would agree with, from a definition that's version-controlled, tested, and auditable. ## What this means for analytics engineers In an agentic world, the analytics engineer's job shifts from competing with AI on code production to defining the semantic layer that AI agents need to be trustworthy. That means metric definitions, contracts, governance ownership, and lineage documentation. It means making the implicit knowledge about what data means explicit enough that a model can reason about it reliably. It means being the person who decides what correct looks like before the agents start running. This is the higher-leverage version of the job. The definitions you write are used by every AI agent that runs against your data stack, not just your team. The governance you put in place is the foundation that makes autonomous pipelines trustworthy. Every architecture diagram circulating today shows AI agents running data pipelines. The ones that work in production will have a transformation layer in the middle, owned by analytics engineers, that defines what the data means and enforces what correct looks like. That layer is dbt. The people who build it matter more now, not less. --- --- title: "Context engineering is the new analytics engineering skill: a practical guide for dbt users" description: "Analytics engineers are already doing context engineering. Here's how your dbt project becomes context for AI." url: "https://www.getdbt.com/blog/context-engineering-is-the-new-analytics-engineering-skill-a-practical-guide-for-dbt-users" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Context engineering is the new analytics engineering skill: a practical guide for dbt users ## What is context engineering? When the idea that 2026 is "the year of context" started making the rounds, Joe Reis answered with a [send-up of context lakes, context products, and the analyst singularity](https://joereis.substack.com/p/gartner-declares-2026-the-year-of). Underneath the satire is a real problem: most organizations are bad at giving AI systems the context they need to reason reliably. Context engineering is the practice of structuring information so that AI models and agents can use it accurately. It works at a different level than prompt engineering, which operates on instructions, and it's broader than retrieval-augmented generation, which is one specific implementation. Context engineering decides what information an AI system gets, in what format, and at what level of specificity, so it can reason about a domain without hallucinating or guessing. For AI systems working with data, that discipline has a specific shape. The AI needs to know: what does this metric mean? What is the grain of this table? What does a null value in this column represent? What business rules make this metric different from that one? Those are questions about how knowledge is encoded in the data layer. And analytics engineers have been encoding that knowledge in dbt projects for years. ## Why analytics engineers already have the advantage Here's the insight this piece is built around: analytics engineers are already doing context engineering. They just don't call it that. When you write a [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definition, you're encoding context. You're making explicit the business logic that turns a column in a table into an answer a stakeholder can trust. When you write a model description that explains why a model exists and what it represents, you're encoding context. When you define a [schema contract](https://docs.getdbt.com/docs/mesh/govern/model-contracts), you're encoding context about what can and cannot change at this model's interface. The difference between an AI agent that reasons reliably about your data and one that hallucinates is almost entirely a function of how well-structured that context is. Agents that query raw tables with no semantic definitions make reasonable guesses and are frequently wrong. Agents that query a governed [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) return answers that are consistent and defensible. The skill analytics engineers already have, making implicit business knowledge explicit, structured, and machine-readable, is exactly the skill that makes AI systems trustworthy. Context engineering is analytics engineering pointed at a new consumer. ## How dbt structures context for AI agents A well-structured dbt project provides context to AI systems through four mechanisms. **Model descriptions.** A description field in a model's YAML does more than document for humans. When an AI agent reads the lineage graph through the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), model descriptions are part of the context it receives. A sparse or absent description means the agent has to guess what the model does. A precise one means it doesn't. **MetricFlow metric definitions.** These are the most powerful context mechanism in the dbt stack. A [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) definition encodes not just what to compute, but how: the entity, the measure, the filters, the grain. When an AI agent queries the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), it isn't writing SQL. It's querying against definitions that already encode the business rules. The accuracy gain is large: in an [early dbt experiment](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms), LLMs answered natural language questions correctly about 83% of the time when grounded in governed MetricFlow definitions, far above raw text-to-SQL on real-world schemas. **Schema contracts.** [Contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) define the interface at a model boundary. They're context about what can be relied upon: which columns will always exist, what types they'll be, what constraints hold. For AI agents writing code that depends on downstream models, contracts prevent a class of errors that would otherwise need human debugging. **Column-level lineage.** Lineage tells an AI agent where data comes from. When an agent generates documentation or debugs an anomaly, [column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) in dbt Explorer is the context that lets it reason about causality rather than just correlation. ## Three context engineering patterns to build this week Three patterns analytics engineers can implement using dbt this week: **Pattern 1: Metric-as-contract documentation** For each of your core business metrics, write a [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) definition that includes the measure plus explicit documentation of what the metric excludes, what the grain means, and what edge cases apply. This is the work that keeps your AI analytics tool from returning a plausible-but-wrong answer to a CFO query. Estimated time: 2-4 hours for 5-10 core metrics. The test: can a new team member, or an AI agent, read the definition and understand not just how to compute the metric but why it's defined this way? If not, the definition isn't doing its job as context. **Pattern 2: Model-level context headers** Add a structured description block to your 10 most important dbt models. Include what the model is (the grain, the purpose), what it isn't (common misuses or misinterpretations), and what changes downstream if its definition changes. This is the context an AI agent needs to reason about the model correctly without running the data. Keep the description complete in 3-5 sentences. Shorter is better if it's precise. The goal is to write everything an AI system needs to reason about the model accurately, no more. **Pattern 3: Freshness and quality metadata as context signals** Configure [source freshness](https://docs.getdbt.com/docs/deploy/source-freshness) checks and make the freshness metadata accessible. An AI agent that knows a source was last loaded 4 hours ago can make a better decision about whether to run a downstream pipeline. An agent with no freshness context runs regardless. Freshness metadata is context engineering at the pipeline-operations level: it gives autonomous systems the information they need to decide when to run. ## The vocabulary shift that matters Context engineering is a better name for work analytics engineers have always done: making implicit business knowledge explicit, structured, and reliable. The consumers of that work have changed. The work itself has held steady. What's changed is the leverage. In 2023, a well-documented dbt project made a human analyst's job easier. In 2026, a well-documented dbt project makes every AI agent that runs against your data stack more accurate. The context an analytics engineer builds is used by a system. That's a meaningful expansion of impact. The year of context, whatever the analysts mean by it, is really the year analytics engineers discover that the work they've always done is worth more than they realized. --- --- title: "Building the agentic data stack: A practical dbt guide for the AI era" description: "AI builds data infrastructure fast. Here's how to make your dbt project ready to support agents without it falling over." url: "https://www.getdbt.com/blog/building-the-agentic-data-stack-a-practical-dbt-guide-for-the-ai-era" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # Building the agentic data stack: A practical dbt guide for the AI era ## The AI readiness gap is bigger than it looks Ali Ghodsi [said earlier this year](https://www.cnbc.com/2026/02/09/under-the-hood-of-the-ai-economy-with-databricks-ceo-ali-ghodsi.html) that over 80% of databases on Databricks' platform are now built by AI agents. AI-driven data engineering has arrived. The headline left out the harder truth. Most of those databases run on foundations built for human-operated workflows, not for agents. Only 7% of enterprises say their data is completely ready for AI, according to the [2026 Data Readiness Index from Cloudera and Harvard Business Review Analytic Services](https://www.cloudera.com/about/news-and-blogs/press-releases/2026-03-05-only-7-percent-of-enterprises-say-their-data-is-completely-ready-for-ai-according-to-new-report-from-cloudera-and-harvard-business-review-analytic-services-reveals.html). The tests, contracts, semantic definitions, and lineage that make AI-generated data trustworthy are still the exception. Both facts describe the same gap. AI builds fast, and without a governed foundation underneath it, fast can also mean unreliable. Accuracy is where it shows: agents that look strong in a pilot slip the moment they meet real users and real edge cases. This gap is exactly what the June 1 launches set out to close. Fivetran and dbt Labs are [one company now](https://www.getdbt.com/blog/what-we-announced-at-snowflake-summit-and-why-it-matters), focused on open data infrastructure for trusted agents. dbt State, dbt Wizard, and dbt Core v2.0 all point at the same goal: a data foundation that's open, portable, and ready for what agents need. That leaves you with one practical question. What does your dbt project need to support AI agents safely? Four things, working together. ## Trust foundation: tests and contracts Start here. [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) enforce schema constraints at the model level, so breaking changes get caught before they reach production. [Data tests](https://docs.getdbt.com/docs/build/data-tests) validate the data itself. For AI-generated models, both matter more. An agent does not know which downstream model breaks when a column type changes. Contracts and tests catch what the agent never thought to check. ## Query reliability: the dbt Semantic Layer Without a semantic layer, agents write raw SQL against undecorated tables, and accuracy drops sharply. In an [early dbt experiment](https://www.getdbt.com/blog/semantic-layer-as-the-data-interface-for-llms), LLMs answered natural language questions correctly about 83% of the time when grounded in governed [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) definitions, far above the accuracy of raw text-to-SQL on real-world schemas. The [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) turns a natural language question into a consistent, defensible answer. Without it, your agent guesses what your metrics mean, and it guesses with confidence. ## The agent built for the work: dbt Wizard General coding agents write plausible code, then hand you an hour of verification because they do not know your project. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) is built for the way analytics engineers actually work: investigating, building, validating, and shipping. It is grounded in your dbt project, so it knows your lineage, contracts, tests, and metric definitions before it writes a line. It knows which tool to call, validates its own work, and shows you what changed and why. You get it inside the dbt platform and as a terminal-native CLI, whether your team runs on the dbt platform or self-hosted. Under the hood, dbt Wizard draws on [dbt Agent Skills](https://docs.getdbt.com/blog/dbt-agent-skills), the open-source skills maintained by dbt Labs and the dbt community. ## Efficiency at AI speed: dbt State Point an agent at your pipelines and it will run them at machine speed. Rebuild everything on every run and your compute bill climbs fast. [dbt State](https://www.getdbt.com/product/dbt-state) checks your metadata and model SQL on each run, builds what's changed, and skips what hasn't, for an average of 30% reduction in warehouse compute. It runs in the dbt platform and in the orchestrator you already use. Build what's changed, skip what hasn't. For agentic workflows, that is the difference between automation that scales and a compute bill that punishes you for using it. ## Audit your own project Before you invest in new tooling, run this audit on the project you have. Each question maps to one of the four pillars. **Trust (contracts and tests)** - Are [model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) enforced on your most important models? If your revenue and customer models have no contract, an AI-generated change upstream can break them quietly. - Do your critical models have [data tests](https://docs.getdbt.com/docs/build/data-tests), and do breaking changes fail in CI? If a schema change passes CI without failing, your governance is misconfigured. **Query reliability (Semantic Layer)** - Are your [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metrics defined in the project? If your core metrics live only in a BI tool, agents cannot query them in a governed way. - Can an external system reach your semantic layer through the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp)? If not, natural language queries are hitting raw tables. **The agent (dbt Wizard)** - Is your team using a dbt-native agent like [dbt Wizard](https://www.getdbt.com/product/dbt-wizard), or general coding agents that do not understand your project? Generic agents do not respect your contracts or see what breaks three models downstream. - Do agents reach your project through a governed, auditable interface, dbt Wizard or the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), rather than raw warehouse credentials? **Efficiency (dbt State)** - Does every run rebuild everything, including models that have not changed? [dbt State](https://www.getdbt.com/product/dbt-state) skips the work that hasn't changed. - Before you connect agents that trigger runs, do you have cost controls in place? Machine speed and a per-query bill are a costly surprise to discover later. ## What an agentic dbt workflow looks like in practice A prompt or an automated trigger starts a task: write a new model, optimize an existing one, generate documentation. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) reads the lineage graph, works out what the affected models do, and drafts the change, validating its own work against your [contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) and [tests](https://docs.getdbt.com/docs/build/data-tests) before you see a diff. The PR runs through CI with contract checks. If it breaks a downstream contract, CI fails, and the agent either fixes the issue or routes it to a human. When the pipeline runs, [dbt State](https://www.getdbt.com/product/dbt-state) builds only what changed, so the run stays cheap. Merged work feeds updated [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metrics and documentation. The next time an analyst asks a question in plain language, the [semantic layer](https://www.getdbt.com/product/semantic-layer) answers from the new model and cites its lineage. Every step depends on at least one pillar. Remove one and the workflow breaks in a predictable place. ## Common failure modes in agentic dbt workflows Teams that build agentic workflows without these foundations hit the same failure modes repeatedly. **AI-generated models with no tests.** The model runs. The numbers look plausible. Three months later a column gets renamed upstream, the model starts returning nulls, and nobody notices until finance flags a wrong number. [Data tests](https://docs.getdbt.com/docs/build/data-tests) catch this. Untested AI-generated models are technical debt at AI speed. **Plain-language queries against raw tables.** An agent queries a table called `orders`. It cannot tell placed orders from fulfilled or cancelled ones, so it guesses, and the answer is wrong. A governed [semantic layer](https://www.getdbt.com/product/semantic-layer) removes the entire category of failure. **Agents with direct warehouse access.** An agent holding database credentials is a liability. [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) and the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) give agents scoped, auditable, governed access to your project. Use them. **Rebuilding everything on every run.** Connect an agent that triggers pipelines and your compute bill climbs while you are looking elsewhere. [dbt State](https://www.getdbt.com/product/dbt-state) pays only for new work. **Treating AI-generated documentation as final.** AI-generated documentation is fast and usually close, and it is not authoritative. Build a human review step into your workflow before AI-generated descriptions publish. ## Build trust before you build automation Sequence matters. Contracts and tests first, the trust foundation. Then the semantic layer, for query reliability. Then the agent and the cost controls, dbt Wizard and dbt State. Teams that skip the foundation and jump straight to agents automate on ground they cannot trust, and all that earns them is faster data debt. The companies reaching production-grade agentic data workflows are the ones that build the right infrastructure before they connect AI to it. --- --- title: "The trust-speed paradox: Governing AI-accelerated data work" description: "72% of data teams use AI to write code. Only 24% invest in checking what it produces. Here's how to close that gap" url: "https://www.getdbt.com/blog/the-trust-speed-paradox-governing-ai-accelerated-data-work" date: "2026-06-16" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # The trust-speed paradox: Governing AI-accelerated data work If you lead a data team or build inside one, the [2026 State of Analytics Engineering report](https://www.getdbt.com/resources/state-of-analytics-engineering-2026) has a finding worth sitting with. Data trust is now the top priority for 83% of data teams, up from 66% a year ago. At the same time, 72% say AI-assisted coding is part of how they work. Only 24% say the same about AI-assisted observability. Read those three numbers together. Teams are using AI to produce more data, faster. They are not keeping pace on the systems that test, validate, and govern what AI produces. That gap is what happens when you adopt AI acceleration without investing in AI-level quality control. Gartner [predicted in early 2025](https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk) that through 2026, organizations would abandon 60% of AI projects built on data that wasn't AI-ready. That's tracking. The failure mode is always the same. AI gets deployed on infrastructure built for dashboards, not for reasoning. The code ships fast. The data it produces is inconsistent, untested, and undocumented. Stakeholders stop trusting it. The project stalls. That's the trust-speed paradox. AI makes data work faster, and faster data work without governance destroys the trust that made the work worth doing. ## Why the gap is worse than it looks AI doesn't just speed up good data practices. It amplifies whatever practices you already have. A team that runs tests, enforces contracts, and documents its models gets faster and sharper with AI. A team that doesn't gets faster at shipping undocumented, untested data. A distribution shift that used to cause a minor dashboard error becomes a hallucination once an agent is reasoning over the data. A schema change that used to break one report now breaks an entire agentic workflow, silently. The data debt you've carried for a decade turns into a different kind of liability the moment an agent has to trust the answers it's getting. You can't govern your way out of that after the fact. The governance has to live in the infrastructure before the AI runs on top of it. ## What a governed AI workflow actually looks like For practitioners, this isn't a new process. It's the dbt workflow you already know, applied consistently, before you add AI acceleration. Tests before merges. Every model gets [`not_null`, `unique`, and `accepted_values` tests](https://docs.getdbt.com/docs/build/data-tests) at a minimum. [Model contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) are enforced on your critical models. Breaking changes get caught in CI, not by someone eyeballing a dashboard after the fact. Contracts at ingestion. Source freshness checks run. Sources and exposures have owners. When something shifts upstream, your dbt project knows, and the failure surfaces where it should. The semantic layer as the source of truth. Your [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow) metric definitions live in code, not in a BI tool's proprietary layer. A natural language query from an agent returns the same number your CFO uses. The definition is version-controlled, auditable, and consistent everywhere it's queried. That's the whole point of the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl). None of this is a new architecture. It's the disciplined use of what dbt was built for, extended to cover the failure modes AI introduces. ## Five investments that close the trust-speed gap If you're a data leader staring at the gap between how fast your team adopts AI and how ready your governance is, the path forward is a sequence of investments, not a transformation program. **Enforce contracts on your most important models.** Not every model needs one today. Start with the models feeding revenue metrics, customer-facing dashboards, and any AI pipeline. [Contracts](https://docs.getdbt.com/docs/mesh/govern/model-contracts) catch breaking changes automatically, and they document the interface other teams depend on. **Turn on column-level lineage.** If you can't trace a number back to its source, you can't debug an agent that returns the wrong answer. [Column-level lineage](https://docs.getdbt.com/docs/explore/column-level-lineage) gives you that trace automatically in dbt. **Define your core metrics in MetricFlow.** Revenue, active users, churn. If those definitions live in a BI tool's proprietary layer, an agent can't reach them in a governed way. Moving them into the [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) makes them machine-readable and consistent. **Add AI-assisted observability to match your AI-assisted coding.** The 72% to 24% gap is the most fixable part of this. If your team uses AI to write models, invest the same way in monitoring what those models produce. **Assign ownership before something breaks.** Ownership is cheap to set and expensive to reconstruct after an incident. Give every source and exposure a named owner. It's the governance investment that pays back fastest when an AI-generated model breaks and nobody can diagnose it. ## Why dbt is both the accelerator and the guardrail Databricks CEO Ali Ghodsi [told CNBC in February 2026](https://www.cnbc.com/2026/02/09/under-the-hood-of-the-ai-economy-with-databricks-ceo-ali-ghodsi.html) that over 80% of the databases on Databricks' platform are now built by AI agents. AI is building data infrastructure faster than people can. dbt sits in an unusual spot here. It's the tool that makes AI-assisted development faster, through [dbt Wizard](https://www.getdbt.com/product/dbt-wizard) and the [dbt MCP server](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) that ground agents in your dbt context. It's also the tool that makes AI-generated data trustworthy, through tests, contracts, column-level lineage, and a semantic layer that grounds agent queries in governed definitions. Most tools in the stack do one or the other. dbt does both. It's a function of what dbt stores: semantic context in code, version-controlled and machine-readable, at the layer where AI needs to understand what the data means. The teams that close this gap first will be the ones that build the foundation before AI runs on top of it. --- --- title: "Your AI isn't broken. Your data model is." description: "AI proof of concepts work. Production doesn't. The gap between those two things lives in your data model, not your model." url: "https://www.getdbt.com/blog/your-ai-isn-t-broken-your-data-model-is" date: "2026-06-08" authors: ["Dustin Dorsey"] categories: ["Insights"] --- # Your AI isn't broken. Your data model is. _This is a guest post from Dustin Dorsey, Senior Director of Data Engineering at phData. Dustin is a co-author of Unlocking dbt. He works with enterprise teams to build the data foundations that make AI reliable at scale._ There is a pattern showing up in organizations everywhere right now, and it is so consistent it almost feels scripted. A team runs a proof of concept. The AI performs brilliantly. Questions get answered in seconds that used to take days. Executives are impressed. Someone uses the word "transformative." Budget gets approved, the rollout begins, and within a few months the same executives are quietly wondering why the thing keeps giving different answers to the same questions. The data team gets called into a meeting. Someone suggests they need a better model. Someone else suggests better prompts. A third person wonders aloud if they should evaluate a different vendor Almost nobody asks the question that actually matters: Is the problem in the AI, or is it in the data the AI is trying to reason over? In most cases, it is the data. And more specifically, it is not a data quality problem. It is a data design problem. ## Why the POC always works The proof of concept works because it was designed to work. This is not cynicism. It is just how POCs operate. When you stand up an AI pilot, you pick a domain you understand well. You choose datasets that have been curated and used repeatedly. You scope the questions narrowly enough that the answers are unambiguous. You have subject-matter experts in the room who course-correct when something looks off. And critically, you are working within a slice of your data where meaning has already been implicitly enforced through months or years of human use. The AI is not discovering meaning in those scenarios. It is operating inside a perimeter where meaning was already established, and it is doing a very good job of navigating that perimeter quickly. The problem is that this creates the impression that your organization is ready for AI at scale. It is not. It is ready for AI in the narrow, well-maintained corner of your data estate that you chose for the demo. The rest of your data is a different story. When AI moves into production, the perimeter disappears. Real users ask questions that span domains. They phrase things differently. They ask about concepts that exist in three tables with slightly different definitions. They want to compare metrics that two separate teams calculate in two separate ways, and both of them call it the same thing. The AI, without a human expert in the room to catch the ambiguity, picks an interpretation and runs with it. Sometimes it picks correctly. Often it does not. And the frustrating part is that you cannot always tell which is which from looking at the output. This is where confidence erodes. Not because the technology failed. Because the expectations were built on a foundation that was never as solid as the demo made it appear. ## The human buffer no one talks about Here is the uncomfortable truth that most AI discussions sidestep: your analysts have always been compensating for this problem. Every time a business user asks "what was our revenue last quarter," an experienced analyst does not just run a query and send back a number. They instinctively clarify intent. They know that "revenue" means three different things depending on who is asking. They know the CFO wants recognized revenue on the accrual schedule; the sales team wants booked orders net of cancellations; and the operations team wants billed invoices for the period. They know which dataset is authoritative for each use case. They know which edge cases to handle and which filters to apply. They run the query, sanity-check the result against a number they already have a rough expectation for, and only then do they send it. That entire process is invisible to the business. It looks like "the analyst ran a query." What it actually is, is a highly experienced person serving as a translation layer between a messy, ambiguous data environment and a decision that needs to be made. AI does not have that translation layer. It cannot. It sees the raw structure of your data, reasons over what it finds, and produces an answer. If the structure is ambiguous, the answer will be inconsistent. Not randomly inconsistent, which would at least be easy to catch, but defensibly inconsistent. Inconsistent in ways where every answer it gives is technically justifiable based on what the data says. That is the hardest kind of wrong to catch, because nothing looks broken. The query ran. The numbers came back. The dashboard loaded. The output just happens to be answering a slightly different version of the question than the one the business was asking. ## One question, five defensible answers Let me make this concrete. "What was our revenue last quarter?" In a typical enterprise data environment, this question has multiple technically valid answers. Revenue might exist at the transaction level, the invoice level, or the recognition level. It might include or exclude returns, internal transfers, or credits depending on who configured the pipeline and when. There might be a table in the CRM that tracks bookings, a separate table in the ERP that tracks invoices, and a reconciliation table in the finance system that is the authoritative source for period-close reporting. All of them have a revenue column. All of them have a date field. All of them will give you a number. A human analyst knows which one to use. They know it because they were told, or because they learned it the hard way, or because they asked and someone explained it in a meeting two years ago that was never documented. That knowledge lives in their head, not in the data. Now ask an AI to answer the same question across all of those tables. It will pick one interpretation based on the structure it can see, the column names it recognizes, and whatever contextual signals exist in the prompt. Ask the same question worded slightly differently and it may pick a different interpretation. Ask it twice on different days and you may get two numbers that cannot be reconciled without knowing exactly which path each query took through your schema. The AI is not making mistakes. It is doing exactly what you would do if you were handed a schema with no documentation and asked to answer a business question as fast as possible. It is making reasonable inferences. The problem is that reasonable inferences are not the same as business-defined answers, and at scale, the gap between those two things becomes very expensive. ## Centralized storage is not the same as centralized meaning Most organizations spent the last decade centralizing data. They moved from distributed data marts to cloud warehouses. They built pipelines. They established governance frameworks. They invested in tooling. By most measures of data infrastructure maturity, they are in a strong position. What they did not centralize is meaning. Centralizing storage answers the question of where data lives. Centralizing meaning answers the question of what that data represents and how it should be used. These are completely different problems, and solving the first one does not automatically solve the second. When you bring data from multiple source systems into a single warehouse without establishing a shared interpretation of that data, you have not created clarity. You have created a larger surface area for ambiguity. You have more tables, more join paths, more definitions of the same concept, and more ways for a system trying to reason over that data to arrive at a different answer than the one you expected. In a human-driven analytics environment, this ambiguity gets resolved through people. It gets resolved in the meeting where two teams present conflicting numbers and someone explains which calculation is correct for this context. It gets resolved through tribal knowledge that experienced analysts carry around and apply every time they touch a dataset. It gets resolved through the dashboard filter that is always set to "exclude refunds" even though there is nothing in the underlying table that enforces that rule. AI cannot attend those meetings. It cannot acquire that tribal knowledge. It cannot apply that filter unless someone has encoded it into the data structure itself. The ambiguity that humans have learned to work around for years does not disappear when AI arrives. It becomes visible. It becomes consequential. And it becomes your most important data problem, regardless of how good your LLM is. ## What actually needs to change There is a version of this problem that gets solved by better prompts. If the ambiguity is narrow and well-understood, you can often describe it in the prompt and get consistent outputs. But that approach has a ceiling. You cannot prompt your way to consistency across a data estate where meaning is systematically implicit. At some point, the only real fix is to encode the meaning into the data itself. This is what dimensional modeling is actually for, and it is why people who have been doing data engineering for a long time have been saying for years that the fundamentals still matter. Dimensional models are not a legacy pattern for old-school BI tools. They are the structural mechanism for making business meaning explicit. They organize data around business processes rather than source systems. They separate what happened (facts) from the context required to understand it (dimensions). They declare grain. They make relationships intentional rather than inferred. The reason structure works where documentation and prompts cannot is worth stating directly. A documented definition of revenue gets ignored by the analyst who never found the Confluence page, and it is invisible to the AI that never reads documentation at all. A prompt can describe which table to use, but only for the one query you thought to write the prompt for. A modeled definition of revenue is enforced at query time, for every query, automatically, whether or not anyone remembers the rule exists. AI cannot read intent. It can only navigate structure. That is why the fix lives in the data layer, not in the prompt layer. When data is modeled this way, the questions AI can reliably answer expand dramatically. Not because the model becomes smarter, but because it has fewer opportunities to be wrong. The structure communicates intent. The interpretation is constrained. The answer space is bounded in ways that align with how the business actually defines its processes. This is not about going back to a rigid schema that cannot accommodate modern analytical needs. It is about recognizing that flexibility without structure is not an asset when AI is doing the reasoning. AI needs guardrails in the data layer that tell it what things mean and how they relate to each other. Dimensional modeling provides those guardrails. ## A question worth asking before you buy anything else Before you evaluate a new model, hire a prompt engineer, or stand up another AI platform, ask your team one question: Can you point to a single authoritative dataset for each of your core business processes? Not a general answer. A specific one. If someone asks "which table is the source of truth for customer lifetime value," is there an answer that everyone agrees on? If someone asks "how is an active customer defined," does the data enforce that definition, or does it live in someone's head and get applied inconsistently? If you cannot answer those questions with confidence, the problem is not your AI. It is the foundation the AI is trying to reason over. And no amount of model tuning or prompt engineering is going to fix a foundation that was never designed to communicate business meaning in the first place. The good news is this is a solvable problem. Organizations that invest in getting the foundation right do not just get better AI results. They get better analytics, better reporting, and better alignment across teams. The AI becomes an accelerant rather than an amplifier of existing confusion. The organizations that skip this step and keep tuning the model instead of the data will find themselves in the same meeting six months from now, still trying to explain why the numbers do not match. And that meeting (the one where two teams present the same metric with different values and neither of them is technically wrong) is worth understanding in its own right. Because that scenario is not an edge case. It is the default state of most enterprise data environments, and AI is about to make it impossible to ignore. If you want to go deeper on the structural conditions that need to exist before AI can operate reliably on your data, I've written a full white paper on this topic. [Building the Foundational Layer for Reliable AI on Structured Data ](https://www.phdata.io/offers/ai-data-foundation-whitepaper/)covers why dimensional modeling functions as trust infrastructure, what it actually means for data to be AI-ready, and why ‌organizations that skip this foundation keep struggling in the same ways. ## Where dbt and phData fit This is exactly why phData and dbt fit so naturally together. dbt provides the implementation home for the kind of intentional, process-oriented data modeling this argument is built on, with model-layer structure, testing, documentation, and the dbt Semantic Layer giving teams a practical way to encode business meaning directly into the transformation layer. phData’s role is to help teams operationalize that in practice: aligning on definitions, designing models around real business processes, and turning the principles described here into systems that can actually be built, governed, and trusted in production. --- --- title: "What is enterprise data infrastructure?" description: "Why your GenAI projects need a robust enterprise data infrastructure that acts as a single control plane for your data." url: "https://www.getdbt.com/blog/enterprise-data-infrastructure" date: "2026-06-04" authors: ["Joey Gault"] categories: ["Pulse"] --- # What is enterprise data infrastructure? Data infrastructure isn't anything new. Companies across various industries have striven to create reliable data pipelines that run at the optimal cost. It's one thing to run a couple of pipelines, however. It's another to process data at the scale, speed, and quality required of modern data-hungry applications. This is especially true in our new [Generative AI (GenAI)](https://www.techtarget.com/searchenterpriseai/definition/generative-AI) era. GenAI use cases require a large volume of highly accurate data from across the organization to succeed. A lack of such data is what will keep [up to 30% of GenAI prototypes from making it into production](https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025). Supplying the data needed for GenAI use cases requires scalability and a robust enterprise data infrastructure. In this article, we'll examine what an enterprise data infrastructure is, where companies struggle when it comes to implementing it, and how to build out an enterprise data infrastructure framework that minimizes the heavy lifting. ## What is enterprise data infrastructure? [Data infrastructure](https://www.getdbt.com/blog/data-infrastructure) is the set of systems and processes that businesses use for data management. In the past, many businesses didn't approach data infrastructure in a uniform manner. Instead, they built one-off data pipelines that were hard to manage and maintain. Data infrastructure establishes a technological and business framework that makes data processing uniform, repeatable, and reliable. Enterprise data infrastructure turns the volume up to 11, enabling a scalable framework for data ingestion, transformation, and data analytics that can feed a company's most data-hungry scenarios. Enterprise data infrastructure also prioritizes data quality and compliance. It ensures that data is only accessible by authorized personnel and processed in accordance with all applicable regulatory requirements and industry standards, including regulations such as GDPR and HIPAA. ### Common components of enterprise data infrastructure Enterprise data infrastructure consists traditionally of multiple tiers or components. The key ones include: Data ingestion. Data can come in many types of data and formats: unstructured, structured, and semi-structured. It can arrive from many data sources. The ingestion layer uses systems such as [Fivetran](https://www.fivetran.com/) to import data from sources and stores it in a consistent format for consumption by other downstream applications. This layer often relies on [Extract, Transform, and Load (ETL) or Extract, Load, and Transform (ELT) pipelines](https://www.getdbt.com/blog/etl-vs-elt) to standardize and load data into downstream systems. Data storage. Cloud computing brought us scalable data warehouse systems like [Snowflake](https://snowflake.com/) and [Amazon Redshift](https://aws.amazon.com/redshift/) that can easily store terabytes of structured data on demand. Other cloud-based data storage architectures, such as [data lakes](https://www.getdbt.com/discover/understanding-data-lakes) and data lakehouses, enable storing datasets in unstructured, semi-structured, or hybrid formats using low-cost cloud storage solutions. Data transformation and modeling. A data transformation layer takes data in its raw state and makes it available to derive insights. Data is generally stored in its raw format and then transformed multiple times to serve specific use cases. This layer typically involves tools to model data as well as [to test data models](https://docs.getdbt.com/docs/build/data-tests) to ensure validity and data quality. Data serving. The data serving layer makes data assets available for discovery and use in business intelligence and analytics, where they take their final form and can be used to drive valuable business insights and data-driven decision-making. This layer consists of tools such as data catalogs, visual dashboards, and automated report generation services to deliver a regular stream of insights in a timely manner. Data governance. [Data governance](https://www.getdbt.com/blog/data-governance) is a set of pillars that include data quality, data stewardship, data security, and data management. Rather than being a specific tool, it's a set of processes built into the fabric of the system, supporting access controls and compliance at every stage of the [analytics development lifecycle](https://www.getdbt.com/resources/the-analytics-development-lifecycle). ## How enterprise data infrastructure can hobble or help your GenAI projects As the AI age accelerates, enterprise data infrastructure is more important than ever. Gone are the days when companies could "just get by" with hard-coded, brittle data pipelines that require constant manual intervention. Successful GenAI projects rest upon a data infrastructure that supports automation, is robust, and can scale to manage terabytes and petabytes of data. What's stopping companies from getting there? Recently, [Fivetran created a global benchmark of enterprise data infrastructure spend](https://www.fivetran.com/blog/the-enterprise-data-infrastructure-benchmark-report-2026) to find out where the gaps are between AI expectations and reality. The report found three glaring issues: Operational efficiency is killing AI data scalability. Data teams dedicate 53% of their time to maintenance. The costs involved in data pipeline downtime and repair may cost companies more than $36 million a year. ROI on data integration is low despite massive spend. This manual, high-intervention approach to data pipelines is killing ROI. On average, enterprises allocate 14% of their data budgets to integration. However, only 27% say that ROI exceeded their expectations. Pipeline modernization delivers higher ROI. The good news is that these problems are avoidable. Fivetran found that 47% of organizations running fully managed and standardized data pipelines exceeded their ROI expectations. The decrease in maintenance work enabled their data teams to address more pressing issues in analytics, AI, and governance, thereby increasing data quality and velocity. ## Why enterprise data infrastructure remains a hard problem The thing is, companies don't end up where they are by accident. There's a reason why organizations end up with a hodgepodge of data pipelines written in different languages and running on their own infrastructure, whether cloud-based or on-premises: Data is fragmented. Without a centralized data catalog to track what's available, [a lot of data ends up in silos](https://www.ibm.com/think/topics/data-silos). These data silos use their own formatting and storage conventions, making them difficult to integrate with other corporate data systems. They often go ungoverned and unmonitored, creating a data security risk for the company, including potential data breaches involving sensitive data. Data exists in heterogeneous systems. Data storage systems have evolved over the past several decades. New architectures, such as data lakehouses, and new storage formats, such as [Apache Iceberg](https://iceberg.apache.org/), have brought needed improvements in quality, performance, and governance. This evolution means data is naturally scattered across disparate, heterogeneous data warehouses, data lakes, and other systems. Often, it's not economically feasible or technically advisable to move that data into newer systems. The time, cost, and risk entailed are too high. Building an enterprise data infrastructure takes time and money. With data engineers spending half of their waking hours fighting fires, there's little time left in a month to tackle the bigger issues. Many companies remain stuck in their current data infrastructure because building a replacement from the ground up requires time and computing resources they don't have. ## Creating a single data control plane for analytics and AI Companies looking to support GenAI don't need to centralize all of their data. What they need is an enterprise data infrastructure that supports delivering high-quality data at scale, no matter where it lives. They need adaptability and workflows that can respond to rapidly changing business needs. What they need, in other words, is a data control plane. A [data control plane](https://www.getdbt.com/blog/ai-ready-platform-generative-ai) is an architectural layer that sits over all of your end-to-end data activities. It enables data integration, access controls, data governance, and protection so you can manage the behavior of people and processes in a distributed and dynamic data environment. At dbt Labs, we've long worked to enable this vision of a vendor-agnostic, flexible, and trustworthy data control plane. dbt enables working with your data where it lives, transforming raw inputs into the high-quality data needed for both analytics and AI use cases. This approach gives stakeholders and data teams a real-time view of data health and streamlines the path from raw source to trusted insight. With dbt, you're not starting from zero. dbt provides a rich framework out of the box for enabling a consistent approach to modeling, testing, and publishing data. That means you can focus less on your enterprise data infrastructure and more on your use cases. dbt provides all of the essential tools you need to curate data for analytics and AI: Authoring and testing data transformations locally. Using the [dbt Fusion engine](https://docs.getdbt.com/docs/fusion/about-fusion), data producers can model and test data transformation code locally, using standard tools like Visual Studio Code. Fusion implements a SQL compiler that automatically validates code as engineers write, enabling data producers to work faster and deploy changes more quickly than ever. Version, test, and monitor continuously. All data analytics and AI code is version-controlled and reviewed before going live. Data developers can easily craft [data tests](http://docs.getdbt.com/docs/build/data-tests) to validate their workloads locally. Tests can also be run as part of a [Continuous Integration (CI) pipeline](https://docs.getdbt.com/guides/custom-cicd-pipelines) to verify any changes before deploying to production. This optimizes the development lifecycle and reduces the risk of errors reaching production environments. Document data for users and AI. dbt's built-in support for rich documentation means users can understand where data comes from and what it means. Using the [dbt Model Context Protocol (MCP)](https://docs.getdbt.com/blog/introducing-dbt-mcp-server) server, you can also feed your docs to your GenAI systems as invaluable context for understanding your data. Define global metrics for GenAI apps. Metrics definitions often differ from group to group, causing confusion when the data is fed to GenAI systems. The dbt Semantic Layer provides a single location for defining and accessing key business metrics. You can supply these "gold" metrics to LLMs and GenAI apps to eliminate confusion between competing definitions and improve machine learning model performance. Drive governance across systems. [dbt simplifies data governance](https://www.getdbt.com/product/governance) as well. By standardizing on dbt as a single data control plane, you give everyone in the company a single source of truth for data and metrics. [dbt Catalog](https://www.getdbt.com/product/dbt-catalog) provides a one-stop shop for global data discovery, with data access enforced via role-based access controls (RBAC). Delivering data for AI requires a solid enterprise data infrastructure that regulates access via a centralized control plane. dbt enables building that foundation on top of your existing, heterogeneous data architecture and gives your organization a competitive advantage in a data-driven world. [Contact us today](https://www.getdbt.com/contact) to learn more about how dbt can modernize your approach to data pipelines and set you up for GenAI success. --- --- title: "Building a data stack for trusted AI" description: "Trusted AI requires governed, consistent, and contextual data. Here’s how to build it without tying yourself down." url: "https://www.getdbt.com/blog/data-stack-trusted-ai" date: "2026-06-03" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Building a data stack for trusted AI AI adoption is stalling. Not for lack of ambition, and not because the models aren't capable. The data foundation underneath isn't ready. The numbers back this up. About 16 percent of organizations surveyed have deployed AI agents, according to Gartner. Meanwhile, 80 percent of IT leaders don't believe their data is ready for AI, and over 70 percent worry about governance in a world increasingly run by agents. That last number should be higher. Way higher. Trusted AI requires trusted data: governed, consistent, and contextual. Without it, every AI initiative follows the same arc: a promising pilot, a scaling problem, a stall. The model isn't to blame. The foundation is. ## The shape of the consumer has changed The data stack was built for people. Every layer, from how we model data to how we serve it, assumes a human is sitting at the end of the pipeline: someone with judgment who can pause when a number looks off and decide whether to trust it. That assumption is being overturned. Agents are querying your data continuously, taking action autonomously, and they don't stop to reconcile. They act. The analyst who runs 50 queries a day has been joined by agents that might run 50,000. Where a human analyst can tolerate lag and ambiguity, an agent making inventory, pricing, or routing decisions cannot. The core issue is **context**. An analyst carries context in their head: they know that "customer" in the revenue model means paying subscribers, not trial users, because someone told them that in onboarding. Agents don't have that. They take what they're given. What a column means, what's authoritative versus stale, what's a trusted curated model versus a raw staging table: that context now has to live in the data itself, in metadata, in contracts, in governed definitions that any system can read. ## Why AI is guessing in the dark Here's what actually happens when you point a generative AI model at your databases, warehouses, or BI datasets without a governed context layer: - The AI scans whatever tables and columns it can see. - It guesses which ones to use based on their names. - It pulls data straight from the warehouse, mixing raw, staging, and curated tables alongside siloed SQL in views and notebooks, with no reliable way to know which represents the actual source of truth. The results are consistent. Without governed context, agents: - Generate unreliable SQL because it can't identify the right models - Invent or misapply your metric definitions - Create governance and trust issues with no clear audit trail - Drive up costs as unreliable queries burn tokens and compute All of this happens because the agent is guessing in the dark. The context it needs to behave reliably simply isn't there. ## Two pillars: governance and structured context Trusted AI rests on two pillars. The first is **governance**: control, [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), and quality. The second is **structured context**: the semantics that agents can actually reason over. Without both, every agent reinvents the truth. Over 50,000 companies use dbt in production. The governed structured context layer it provides fixes the context gap. It tells your AI or agent how your data is defined, how it connects, and what it actually means. It exposes the rich metadata that already lives in your dbt project: your models, lineage, metrics, freshness definitions, and documentation. Then, it surfaces all of that through open standards, including the m[odel context protocol (MCP)](https://docs.getdbt.com/docs/dbt-ai/about-mcp) and the [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) using [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow), so any AI system can rely on a single governed source of truth rather than reconstructing context from scratch every time. **[dbt Wizard](https://www.getdbt.com/product/dbt-wizard) **complements this by packaging proven dbt workflows - including [testing](https://docs.getdbt.com/docs/build/data-tests), debugging, migrations, and metrics definition - into reusable patterns. Agents not only know the context; they know how to follow a consistent, proven process when acting on it. **[dbt Wizard CLI](https://docs.getdbt.com/docs/dbt-ai/about-dbt-wizard-cli?version=2.0&name=Fusion) **surfaces that same context for fast, cost-efficient local development. The benefits compound quickly, with structured models and tests giving you AI that generates reliable, reusable SQL. A governed, query-optimized semantic layer gives AI the right definitions and logic, while centralized lineage surfaces clear ownership and makes every change auditable through Git and pull request (PR) workflows. The result is both increased accuracy and greater cost efficiency. When AI works from curated context rather than your entire warehouse, your token, compute, and review costs stay manageable. [**Check out dbt Wizard, your personal dbt agent, available wherever you work.**](https://www.getdbt.com/product/dbt-wizard) ## The semantic layer is not optional The semantic layer is the component that makes AI actually reliable. When metrics and business entities are defined once and reused everywhere, whether the consumer is a dashboard, an operational workflow, or an agent, they all work from the same trusted logic. That consistency is what separates AI that amplifies good decisions from AI that amplifies errors, with no human in the loop to catch the difference. Gartner estimates that by 2027, enterprises without a semantic layer will spend 40 percent more on AI rework and remediation than those with one. Don't be those enterprises. At dbt Labs, we approach this through open standards. [The Open Semantic Interchange (OSI)](https://www.getdbt.com/blog/dbt-labs-affirms-commitment-to-open-semantic-interchange-by-open-sourcing-metricflow) is an open standard for how semantics move between tools. MetricFlow gives you one definition of every metric, consumable anywhere. Define once, use anywhere, and it works with whatever agent or application comes next. The practical upshot: dbt is agent agnostic. We work with whatever LLM framework or vendor you choose. Your governed context travels with the data. As your AI stack evolves with new models, tools, and interfaces, the structured foundation stays the same. You centralize logic in one layer, reduce redundant queries, and get more reliable AI outputs across the board. ## What this looks like at ACV Auctions [ACV Auctions](https://www.acvauctions.com/) runs a wholesale dealer-to-dealer automotive marketplace. Their analytics manager, Darren Peters, has been building through this challenge firsthand. The starting point was familiar. ACV's embedded reporting system couldn't keep pace: simple report tweaks for dealers could take weeks or months. Their internal BI tool had accumulated over 1,500 dashboards. A new employee searching a common business term would see 400 different results, each a variation on the same concept, with no reliable way to identify ground truth. Darren's instinct, well before AI chat was part of anyone's roadmap, was rigorous data governance at the dbt layer. He talks about writing column descriptions the way you'd write a board game rule book: specific enough that no one argues over what a rule means. The goal was traditional data governance: a clean data dictionary, automated quality tests, and descriptions piped from [BigQuery](https://cloud.google.com/bigquery) into [Confluence](https://www.atlassian.com/software/confluence). Precise, unambiguous definitions for every column. That discipline turned out to be the foundation for everything that came after. When ACV brought in [Omni Analytics](https://omni.co/) for their semantic layer, Omni could pull column-level metadata directly out of their data warehouse automatically. Darren built the governance layer for traditional reasons. But it also set up ACV perfectly for AI. A few months into using Omni's AI chat, the team's workflow has shifted. When a [Slack](https://slack.com) message arrives from a stakeholder with a data question, Darren pastes it straight into the chat. If the response looks right, he sends back a shareable query URL: a live, governed analysis the recipient can view or iterate on themselves, without it being formally published and adding to the content pile. No more ad hoc dashboards that bloat the system just to answer a one-off question. Product managers who used to wait days or more for ROI analyses on feature work are now running those analyses themselves. Darren's team validates and adjusts, but a backlog that used to go untouched is actually getting cleared. Customer support teams are answering dealer-specific account questions in real time, without ticket queues. Darren describes himself as a former AI skeptic, specifically in the context of self-serve analytics. Two things shifted his view: the maturity of the semantic layer itself, and access to more capable reasoning models. When both came together, it wasn't a gradual improvement. "It was a light switch," he said. One question in the chat, one response, and he knew it was real. The shift ACV is living now: less time building content, more time engineering context. When the core analytics loop is about defining a dimension precisely and describing it in plain language, the team has to stay close to the business and keep sharpening its understanding. That rigor pays off across every consumer, whether human or agent. ## Shipping AI with confidence Solving for the context gap is what finally lets you ship AI with confidence. It leads to fewer hallucinations, better decisions, lower security and governance risk, reduced token and compute spend, and faster data development. Most importantly, context enables AI initiatives that actually scale beyond the pilot stage. They scale because teams are willing to adopt and trust them. That's the part people underestimate. Governance isn't the enemy of speed. It's the condition for adoption. When data semantics are open and portable, your stack stops being a series of vendor decisions and becomes a foundation you actually own. Every layer, whether ingestion, storage, or semantics, stays interoperable and yours to evolve. That's what lets enterprises scale AI without being held hostage to architectural choices made three years ago. [**For a hands-on look at all of this in practice, watch the full session recording.**](https://www.getdbt.com/resources/webinars/the-future-is-open-building-a-flexible-ai-ready-data-stack-without-lock-in) Build the governed foundation your AI initiatives need: [Talk to the team at dbt today](https://www.getdbt.com/contact). --- --- title: "dbt Labs Named Snowflake Data Integration Product Partner of the Year" description: "dbt Labs wins two Snowflake Partner honors: Data Integration Product Partner of the Year and Snowflake’s CoCo Adoption Award" url: "https://www.getdbt.com/blog/dbt-labs-named-snowflake-data-integration-product-partner-of-the-year" date: "2026-06-02" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Named Snowflake Data Integration Product Partner of the Year _Fourth consecutive annual award honors commitment to exceptional outcomes for joint customers_ **PHILADELPHIA – June 2, 2026: **[dbt Labs](https://www.getdbt.com/), a leader in standards for AI-ready structured data, announced today at [Snowflake Summit 2026](https://www.snowflake.com/summit/) that it has been named the 2026 Data Integration Snowflake Product Partner of the Year, in addition to being recognized for Snowflake’s CoCo Adoption Award for leading adoption and delivering customers transformative results via Snowflake’s coding agent and control plane for builders. dbt Labs is being recognized for its achievements as part of the Snowflake AI Data Cloud, helping joint customers unlock production-grade workflows that are built on a reliable, governed and trusted data foundation, ready to run AI agents reliably, at scale. dbt has become a preferred transformation and context engine for customers’ AI and analytics use cases, with over 75% of customers with Snowflake accounts using dbt and powerful tools like the dbt Semantic Layer, the dbt Fusion engine, and dbt MCP server to unlock a faster developer experience and execute new and complex AI use cases. With 90% of joint customers actively using Snowflake Cortex AI, dbt is an integral part of their AI journey, delivering the reliable data foundation that enables these organizations to take full advantage of Snowflake's AI functionality. “Trust in data is the most widely prioritized organizational objective, and Snowflake Marketplace is a powerful resource to connect enterprises to dbt and its latest features that bring structure, governance, and velocity to what data teams are building in the AI era," said Shawn Toldo, Vice President, WW Partner Organization at dbt Labs. "This award underscores our mutual dedication to supporting our joint customers and delivering remarkable results, allowing for data-driven innovation at scale to expand the reach of data and AI capabilities." This is the fourth consecutive year that Snowflake selected dbt Labs as a Snowflake partner award winner, which is a testament to the depth and durability of their collaboration and the impact of this longstanding partnership. The companies are united in their mission to help customers cost-effectively build AI-powered insights and data assets, ultimately driving deeper organizational trust in data and the teams that power it. With transactions on the Snowflake Marketplace growing by approximately 30% year over year, dbt Labs is helping customers utilize their full spend from Snowflake contract commitments to ensure optimal ROI. "dbt Labs has been an incredible partner to us over the years and we're proud to name them as Snowflake's 2026 Data Integration Partner of the Year," said Amy Kodl, SVP, Worldwide Alliances & Channels. "The work their team is doing with the AI Data Cloud ecosystem continues to deliver strong results for our joint customers." Joint customer [WHOOP](https://www.getdbt.com/case-studies/whoop) is a testament to this continued collaboration. As the WHOOP team grew, they used the dbt platform as a scalable solution to ensure clean and well-governed data was being migrated into Snowflake. dbt provided the foundation that gave all stakeholders at WHOOP access to reliable data, which allowed the team to [use Snowpark](https://www.snowflake.com/en/customers/all-customers/case-study/whoop/) to build the WHOOP AI/ML financial forecasting model. What once was a major roadblock is now streamlined, saving time for the WHOOP analyst and data engineering teams to focus on more strategic initiatives. Learn more about dbt Labs and Snowflake [here](https://www.getdbt.com/data-platforms/snowflake), and visit the dbt Labs booth (#2112) during this week's Snowflake Summit to explore the [latest dbt innovations](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents). Stay on top of the latest news and announcements from dbt Labs on [LinkedIn](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=82576339&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fdbtlabs%2Fmycompany%2F&a=LinkedIn), [X](https://x.com/dbt_labs), [Instagram](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=255235389&u=https%3A%2F%2Fwww.instagram.com%2Fdbt_labs%2F&a=Instagram), and [YouTube](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=508887317&u=https%3A%2F%2Fwww.youtube.com%2Fc%2Fdbt-labs&a=YouTube). **###** **About Fivetran + dbt Labs** Fivetran + dbt Labs deliver the data infrastructure layer that makes agents trustworthy — from the moment data moves, through every transformation, to the context an agent reasons from. The Fivetran platform moves, manages, and transforms data from every system a business runs on into a secure, reliable foundation engineered to evolve, with the flexibility to work across clouds, engines, and tools. With Fivetran, analytics, operations, and AI run on data you trust and control. Thousands of organizations worldwide, including OpenAI, LVMH, Pfizer, and Verizon, rely on Fivetran to turn data into a competitive advantage. Learn more at [Fivetran.com](http://fivetran.com), or follow Fivetran on [LinkedIn](http://linkedin.com/company/fivetran). Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 100,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "What we announced at Snowflake Summit and why it matters" description: "dbt State, dbt Wizard, dbt Core v2.0, and the Fivetran merger" url: "https://www.getdbt.com/blog/what-we-announced-at-snowflake-summit-and-why-it-matters" date: "2026-06-01" authors: ["Corinne Hallander"] categories: ["Product"] --- # What we announced at Snowflake Summit and why it matters We showed up at Snowflake Summit differently this year. In case you missed it: [Fivetran and dbt Labs are now a single company](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents), with a shared vision of building a foundation for trusted agents using open data infrastructure principles. There's a lot to cover among all the announcements, so let's get into it. ## Fivetran and dbt Labs have merged If you've worked in data for the last decade, you know both names. Fivetran is the standard for getting data into your warehouse reliably and automatically, with connectors for virtually everything. dbt is the standard for data transformation. For years, teams have run Fivetran and dbt together. The combination was already the backbone of the modern data stack for thousands of organizations. What changes now is that the Fivetran and dbt teams are working toward the same thing: a data foundation that’s open, portable, and ready for what agents need. Customers of both Fivetran and dbt are already leading successful, trusted AI and agent initiatives, backed by our products. The value of this combined approach was emphasized by Piyush Bhargava, Sr. Director Global Data & Analytics at DocuSign: "AI and agents are only as strong as the data behind them. By investing in Fivetran and dbt, we've built the reusable, trusted data assets that are central to how we scale AI and drive innovation." [You can read a lot more about the merger thesis, dbt's continued commitment to open source, and what's next in Tristan's blog](https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means). ## dbt Core v2.0: The open-source foundation, rebuilt on Fusion dbt Core started in 2016 as a way for data practitioners to work like software engineers with modular SQL, version control, testing, and documentation in an open-source framework. It quickly became the industry standard, and now over 100,000 organizations run it today. But the original dbt engine—built in Python—has its limits. Parse times grew with project size. The execution layer had no understanding of the SQL it was running. Errors surfaced only after hitting the warehouse. A ground-up rewrite was needed. That rewrite became the dbt Fusion engine: built in Rust, with native SQL comprehension and parse times up to 10x faster than the original engine. We launched Fusion last year and today, over 4,500 projects run on Fusion. But maintaining two engines meant confusing licensing, an all-or-nothing-feeling migration path, and real uncertainty for contributors, adapter maintainers, and partners about where to build. So this week, we changed that. We open-sourced the Fusion runtime and released it as dbt Core v2.0 under the Apache 2.0 license. This means that the two-engine era is coming to an end, and every practitioner can enjoy the dbt experience they know on a faster, more scalable foundation. Consolidating to one engine also means that our commercial investments will directly drive improvements in the open-source distribution of dbt. This is the biggest open-source expansion that dbt has made in years. Fusion extends dbt Core v2.0 in a new proprietary distribution with richer capabilities, such as SQL comprehension, column-level lineage, and instant feedback while you work. dbt Core v2.0 is in alpha release, while the proprietary distribution is in preview. The most used adapters are supported at launch: Snowflake, BigQuery, Databricks, and Redshift. The path forward is simple: `pip install dbt==2.0.0-preview.x` for a single binary that brings together the open runtime, Fusion-powered capabilities, and access to platform-connected workflows. [**Read more: dbt Core v2.0 is here**](https://docs.getdbt.com/blog/dbt-core-v2-is-here) ## Announcing dbt State: Build what’s changed, skip what hasn’t Every time a pipeline runs, it rebuilds everything, even tables that haven’t changed. You pay for every one of those compute cycles. Over time, this becomes a significant cost, not just in warehouse spend, but in time engineers spend managing job schedules, building sub-selectors, and orchestration around the inefficiency. Custom, manual workflows mean slower data and slower answers. **dbt State changes this**. On every run, dbt State checks your metadata and model SQL to see what’s changed. When upstream data or code has changed, it builds the model. Otherwise, it skips the build by reusing existing state, cloning existing state, or auto-deferring (in development) to production state. In production, this results in an average of 30% reduction in warehouse compute and prevents breakage. In development, this means faster iterations without fear of costly mistakes. CarGurus Vice President of Data Parag Shah put it simply: > "Before dbt State, every job rebuilt every model in the lineage. Every. Single. Time. Now, with dbt State, dbt checks if source data changed. If it didn't, the model is skipped. For us, that resulted in a 9% compute reduction, 35% fewer models built, and a 15% reduction in Snowflake backfill costs." ```json { "_key": "033b51c767a3", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/Wb-b4BjoVUE" } ``` The impact goes beyond compute savings. When you know you're only paying for new work, you stop rationing runs. Teams that used to schedule refreshes carefully, because more frequent runs felt wasteful, can now run as often as the business needs fresh data, without the fear of an expensive bill at the end of the month. And with orchestration logic living in the project rather than a spreadsheet of job definitions, engineers spend less time on maintenance and more time on what matters. State-aware orchestration has been delivering cost savings to dbt platform users on Fusion during preview. dbt State opens that capability to _every_ dbt user: running locally on dbt Core (1.7+) or the dbt platform, in your orchestrator of choice, true to open data infrastructure principles. [**Available now in Preview: get started today**](https://www.getdbt.com/product/dbt-state). _Want to learn more?_ [Join us for a live virtual event on July 15th to see dbt State in action](https://www.getdbt.com/resources/webinars/dbt-state-build-what-s-changed-skip-what-hasn-t/?utm_medium=internal&utm_source=blog&utm_campaign=q2-2027_dbt-state-deep-dive-product_aw&utm_content=themed-webinar____&utm_term=all_all__) with a real customer story and answers to the questions practitioners always ask before they deploy. Save your seat now. ## Announcing dbt Wizard: an AI agent built for analytics engineering Coding agents are good at writing decent code. They're less good at writing code that knows your dbt project: code that respects your upstream models, doesn't break downstream dependencies, and handles the governance requirements your team has spent years building. The agent generates something plausible, but you spend the next hour verifying it. Analytics engineering is not the same as software engineering. You're not just editing files. You're reasoning about lineage. You're aware of what breaks three models later. You're accountable when something goes wrong in production. **Generic tools weren't built for that context.** **dbt Wizard is.** dbt Wizard is an AI agent built specifically for the analytics engineering workflow: investigating, building, validating, and shipping. It's grounded in your dbt project natively. That means it knows your lineage, your contracts, your tests, and your metric definitions before it writes a single line. It knows which tool to call. It validates its own work. It shows you what changed and why. The difference this makes is concrete. > “Before dbt Wizard, our engineers were spending more time correcting AI output than they were writing models. Now the agent actually knows our project. It gets the joins right, it respects our contracts, and it doesn't break things downstream. We've seen a 15-20% reduction in production incidents since we rolled it out." - Erion Krasniqi, Junior Data Scientist, Endress+Hauser InfoServ ```json { "_key": "cce9e4d9c7d5", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/sQ8VQbOimng" } ``` dbt Wizard is available in two surfaces: inside the dbt platform and from the terminal via dbt Wizard CLI for teams developing locally. It works across Snowflake, BigQuery, Databricks, Redshift, and every other dbt-supported warehouse. [Get started today](https://www.getdbt.com/product/dbt-wizard). _Want to learn more?_ [Join us for a live virtual event on July 22nd to see dbt Wizard in action across the CLI and dbt platform](https://www.getdbt.com/resources/webinars/dbt-wizard-an-agent-purpose-built-for-analytics-engineering/?utm_medium=internal&utm_source=blog&utm_campaign=q2-2027_dbt-wizard-deep-dive_aw&utm_content=themed-webinar____&utm_term=all_all__)—real use cases, real analytics teams using it today, and answers to the questions you'll have after watching the demo. Save your seat now. ## Recognized as Snowflake's Data Integration Partner of the Year The week brought more than product news. At Snowflake Summit, dbt Labs was named the 2026 Data Integration Snowflake Product Partner of the Year, along with Snowflake's CoCo Adoption Award for leading adoption of Cortex Code and delivering customers transformative results. It's the fourth consecutive year Snowflake has recognized dbt Labs with a partner award, a reflection of how deep the collaboration has become. The reason is simple: trusted data is the foundation for trusted AI. Over 75% of customers with Snowflake accounts use dbt, and 90% of joint customers actively use Snowflake Cortex AI. dbt is the preferred transformation and context engine for those AI and analytics use cases. [Read the full announcement](https://www.getdbt.com/blog/dbt-labs-named-snowflake-data-integration-product-partner-of-the-year). ## What’s next? For data teams, the question has always been: how do we do more with what we have? Better pipelines, smarter infrastructure, less time managing things that should manage themselves. That's what we're building. Still have questions? On June 25th, Tristan Handy (President and co-founder, Fivetran & dbt Labs) and Taylor Brown (COO and co-founder, Fivetran & dbt Labs) are hosting a [live virtual event](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a) to walk through the merger, what's shipping in dbt, and what it means for your stack, and then taking your questions live. This is the most direct access you'll have to the people making these decisions. [Save your seat today](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a). --- --- title: "Fivetran and dbt are one company now. Here's what that means." description: "Fivetran and dbt Labs are officially one company to deliver data infrastructure for agents you trust." url: "https://www.getdbt.com/blog/fivetran-and-dbt-are-one-company-now-here-s-what-that-means" date: "2026-06-01" authors: ["Tristan Handy"] categories: ["Company"] --- # Fivetran and dbt are one company now. Here's what that means. We [announced](https://www.getdbt.com/blog/dbt-labs-and-fivetran-merge-announcement) the merger in October. Today it's official: [Fivetran and dbt Labs are one company](https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents), with a shared mission to build open data infrastructure for the agentic era. The timing couldn't be better. In the months since October, the shift toward agentic AI has gone from a curiosity to a force of nature. What started as a strategic combination of two organizations to deliver best-in-class data pipelines has become something even more urgent: an absolute imperative to prepare the world for agents. We've spent years building complementary halves of the same stack. Now we get to build them together…at exactly the moment it matters most. ## The thesis Agents are rapidly becoming the primary consumers of your enterprise data. Agents are no longer just in production at frontier companies, but in enterprises globally. Agents approve insurance claims, deflect support tickets, and write code. And they do it at increasingly massive volumes: up at least 1-2 orders of magnitude YoY. Many of these agents will need access to your organizational data. Unfortunately, most data infrastructure deployed today was not built for agents, but for humans: analysts running queries, dashboards refreshing on a cadence, data consumers clicking around in a BI tool. Humans are patient. Humans bring their own context. Humans do one task at a time. Agents don’t. Agents operate at machine-speed. They’re running 24/7. They can issue queries at a volume human traffic never approached. And crucially: they don't have the institutional knowledge that a human analyst carries in their head. They can't fill in the gaps. When they encounter undefined metrics, ungoverned data, or fragmented business logic, they don’t shoulder-tap their colleagues...they often just (confidently!) produce the wrong answer. And the more autonomous they are, the harder that wrong answer is to identify and troubleshoot. 80% of IT leaders believe their enterprise data is not ready for agentic AI. 70% worry about AI governance with the proliferation of agents, according to Gartner. These aren't anxieties about the future. They're descriptions of the present. The bottleneck has shifted. For the past decade, the bottleneck in data was infrastructure, and we, collectively, largely solved it. The modern data stack worked, although sometimes it could be a bit of a pain. **The new bottleneck is trust**: trust in the data, the context, the governance. That is precisely the problem we built to solve. Together. ## What Fivetran and dbt each bring, and why coming together matters **Reliable, governed, and trusted data for agents.** Fivetran ensures data is complete, fresh, and reliably moved. When an agent reaches for your data, it's current and it reflects the actual state of your business across many sources. Agents querying stale, siloed data don't just produce wrong answers—they take wrong actions. Freshness matters more when the consumer is autonomous. But fresh data alone isn't enough. Agents don't bring context; they inherit it. dbt's role is making sure that context is governed, defined, and trustworthy: business logic versioned in code, metrics defined once and tested, lineage traced from raw source to downstream consumer. This matters well beyond conversational analytics. The agent approving a loan, deflecting a support ticket, or triggering a supply chain reorder needs the same thing an analyst needs: reliable data. When an agent asks, "what’s the shipment status of this order?" or “is this customer in good standing?” it has to know exactly how to get that answer, without guessing. **Flexible and portable. Any engine, cloud, or model. No lock-in.** Here's the thing about that business context: it needs to travel. The AI landscape is shifting fast: new models, compute platforms and tools are coming in and out of favor so fast that organizations can’t fully adopt one framework before the tide turns and it’s onto the next. The platform-native approach to data infrastructure—where your semantic definitions and pipelines are tightly coupled with a single compute vendor—breaks down the moment you need to either migrate or go multi-engine. Fivetran moves data across any source without owning your storage. dbt keeps your business logic in code you control, portable across any engine or cloud. Deep integrations across the AI and analytics ecosystem means this infrastructure is pluggable at every layer, even as you migrate or go multi-engine. Together, you get a flexible, portable foundation that evolves with your architecture instead of constraining it. That's not a feature: it's a core design principle. **Scalable and optimized for production AI demands**. Agents will issue queries at a volume human analysts never approached, and without an architecture built for it, costs grow incredibly fast. The data supports this: Despite a[ 95% drop in token costs](https://venturebeat.com/orchestration/cheaper-tokens-bigger-bills-the-new-math-of-ai-infrastructure), enterprise AI spend has[ exploded from $1.7B in 2023 to $37B in 2025](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/). Cheaper tokens, yet bigger bills. Agents repeatedly querying fragmented systems for missing context is expensive in compute and tokens. Managed data connections reduce operational overhead. Governed context minimizes unnecessary retrieval hops. Decoupled storage and compute means you can route workloads to the most efficient engine for each job. The goal is for per-query cost to fall as volume scales, not rise with it. **Together, Fivetran and dbt deliver open data infrastructure for agents you trust, at scale:** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/58d77b8611ec19dd44538560287d8192b86599e5-2048x1645.jpg) This is just the beginning. Today, we announced [two new exciting dbt features](http://www.fivetran.com/fivetran-dbt-labs-merger) available to users wherever they write code. ## Our commitment to open source, and what we're doing to reinforce it Our commitment to open source isn’t changing. Open collaboration and community contributions have defined dbt's evolution from the start. It’s the very thing that’s made dbt the de facto standard for transformations on structured data, and that same ethos is what’s going to propel our users to succeed in this AI wave. Today, [we announced the alpha of **dbt Core v2.0**](https://docs.getdbt.com/blog/dbt-core-v2-is-here): the next major version of dbt Core. As it has been since the first commit, **dbt Core remains licensed under Apache 2.0 **with this release. What changes in v2.0 is what's under the hood: the kernel of the dbt Fusion engine, a full Rust rewrite of the dbt runtime, is being open-sourced under Apache 2.0. The Rust foundation powering Fusion—faster parsing, more consistent execution, dramatically better performance in agentic workflows—is now available to the entire community under the same license you’ve been using for a decade. Instead of two codebases, two licenses, and ambiguity about what’s available on dbt Core versus Fusion, where to build or how to migrate, **we will simply have one engine: dbt**. Not only are we excited to deliver all of these benefits to every dbt user, having a single engine to invest in allows us and the dbt community to innovate faster in the AI age where speed is everything. This is the single biggest drop of new Apache 2.0-licensed code we've shipped in years…potentially ever. ## What’s next Personally, I’m excited about the future of the data ecosystem, of the role of the data practitioner, of what Fivetran and dbt Labs can build together. After several years of fascinating progress in consumer AI but a constant refrain of it-doesn’t-quite-work-yet in the enterprise, _AI and agents are truly here_. Which has caused the pace of change for everything inside of the data ecosystem to absolutely skyrocket. This is great. This is what we should all want. Data practitioners—from engineers to analysts—are critical to the AI future, but have been sitting on the sidelines for the past few years saying “put me in, coach!” Guess what: **it’s time**. While none of us know _exactly_ how the next several years will play out, things will move quickly, and data practitioners everywhere will be central to the story. What you can see us doing—with the merger, with our product launches, with our investments in open data infrastructure—is to answer the question “How do we position ourselves to elevate and advocate for data practitioners into the next decade?” The “how” is changing, but the mission remains constant. **We want to be your partners in building data infrastructure for the age of AI and agents.** If you have questions about our shared vision, Core v2.0, and new products we announced at Snowflake Summit, [join me on June 25th for a live Q&A](https://www.getdbt.com/resources/webinars/fivetran-dbt-labs-the-merger-what-s-shipping-in-dbt-and-live-q-and-a/?utm_medium=internal&utm_source=blog&utm_campaign=q2-2027_fivetran-dbt-merger_aw&utm_content=themed-webinar____&utm_term=all_all__). --- --- title: "Fivetran + dbt Labs Complete Merger to Create the Data Infrastructure for Trusted AI Agents" description: "Fivetran and dbt Labs are now one company, focused on building the data foundation for the agentic AI era." url: "https://www.getdbt.com/blog/fivetran-dbt-labs-complete-merger-to-create-the-data-infrastructure-for-trusted-ai-agents" date: "2026-06-01" authors: ["Elaine Green"] categories: ["Press"] --- # Fivetran + dbt Labs Complete Merger to Create the Data Infrastructure for Trusted AI Agents **OAKLAND, Calif. — June 1, 2026 — **[Fivetran](http://www.fivetran.com), the data foundation for AI, today announced the completion of its merger with [dbt Labs](http://getdbt.com), the creator of dbt and the leader in standards for AI-ready structured data. Initially operating as Fivetran + dbt Labs, the all-stock transaction, originally [announced](https://www.getdbt.com/blog/dbt-labs-and-fivetran-sign-definitive-agreement-to-merge) on October 13, 2025, brings together two category-defining platforms to advance a new era of trusted, Open Data Infrastructure for AI at scale. George Fraser will continue serving as CEO, and Tristan Handy will serve as President. Together, Fivetran + dbt Labs support a global community of more than 100,000 data teams across analytics, data engineering, and AI initiatives, including some of the world’s most recognized brands such as OpenAI, Zendesk, Coupa, and HubSpot, as well as leading enterprises across financial services, retail, manufacturing, and healthcare. **A new foundation for agentic AI** AI agents are quickly becoming the primary consumers of enterprise data, and they behave differently from the human analysts that today's data stack was built to serve. Agents operate continuously, in parallel, and at machine speed. And many organizations want to move into a world where most agents are autonomous — no human in the loop. This shift raises the bar for the data they run on, requiring it to be reliable, fresh, governed, and accessible across every system in the enterprise. Fivetran + dbt Labs are building the data foundation for the agentic AI era. Together, these companies deliver the data infrastructure layer that makes agents trustworthy, from data movement and transformation to the governed context needed for reasoning and action. Fivetran ensures agents operate on complete, continuously synced, and reliable data. dbt ensures that data is defined, tested, and trusted through governed business logic, shared semantic context, and software engineering best practices embedded throughout the data lifecycle. Built on open standards, this foundation works across any cloud, engine, and tool, giving organizations the freedom to evolve their architecture without lock-in while maintaining portable business logic, cost efficiency, and control. “The next generation of enterprise AI will be defined by the quality and trustworthiness of the underlying data,” said George Fraser, CEO and Co-Founder of Fivetran + dbt Labs. “Together, Fivetran and dbt Labs are creating the infrastructure layer that helps organizations deliver governed, high-quality, and semantically rich data to power trusted AI agents at scale.” "The companies that deploy AI successfully over the next decade will be the ones whose agents can be trusted to act," said Tristan Handy, President and Co-Founder of Fivetran + dbt Labs. "Trust is built at the infrastructure layer, on high-quality tooling and on open standards. That's the bet we're making together." **Joint product innovations** The merger also marks the first major milestone in a shared innovation roadmap, with the first combined innovations from Fivetran + dbt Labs debuting today. The announcements span agentic development workflows, intelligent orchestration, and continued investment in open source innovation, including extending powerful dbt capabilities to the dbt Core open source user base. Key innovations announced today include: - **dbt Core v2.0 (alpha): **The open sourcing of the dbt Fusion engine runtime, released as dbt Core v2.0 under an Apache 2.0 license, giving every practitioner the dbt experience they know on a faster, more capable foundation. Additionally, the locally installable distribution of dbt gives developers free access to the full breadth of Fusion’s capabilities — core language features and warehouse adapters — with the ability to seamlessly unlock additional platform features by logging in directly from the terminal. - **dbt State (preview):** dbt State acts as a caching layer for data pipelines. It only builds what’s changed and skips what hasn’t, helping companies reduce underlying infrastructure costs by 30% or more. - **dbt Wizard (beta): **dbt Wizard brings autonomous assistance for model authoring, refactoring, and debugging, grounded in full dbt project context, including lineage, tests, contracts, and defined metrics. The result is governed recommendations and trusted SQL generation that reflect how enterprise data is actually structured and defined. - **Agents Schema:** An open source standard for agentic context that designates a single schema in the warehouse or lake as the shared context layer for AI agents. Metric definitions, semantic models, dbt lineage, and business documentation are stored in plain SQL tables and can be published from existing systems through tools such as GitHub Actions, metadata connectors, or custom integrations. Compatible with any warehouse, lake, ingestion tool, or SQL-capable agent, Agents Schema gives organizations a customer-owned context layer that works within existing security and governance policies, improves token efficiency through richer context, and eliminates the need for new infrastructure or vendor-locked agent systems. **Customers building with Fivetran + dbt Labs** “With Fivetran and dbt, what used to take months now happens in weeks,” said Akshay Agrawal, Director of Data Engineering at Zendesk. “It gives the business faster access to trusted data and creates the foundation we need to scale analytics, agents, and AI across the enterprise.” “Our focus now is on how we operationalize AI across Inova. With Fivetran and dbt, we’re creating the foundation for AI agents and applications that can act on trusted, governed data — not just generate insights, but drive action,” said Jon McManus, Chief Data and AI Officer, Inova Health. “The combination of Fivetran and dbt isn't just about efficiency today,” said Lakshmi Ramesh, VP of Data Services at Tinuiti. “It's about being ready for what's next. Analytics, AI, and agentic workflows all run on trusted data, and together, Fivetran and dbt are the data infrastructure that makes it all possible.” “By building an AI-ready data foundation with Fivetran and dbt, we’re improving how teams across Shutterstock access and operationalize trusted, real-time data for analytics and emerging AI-driven workflows,” said Jitesh Kumar, Senior Software Development Manager at Shutterstock. "AI and agents are only as strong as the data behind them,” said Piyush Bhargava, Sr. Director, Data Architecture and Engineering at DocuSign. “By investing in Fivetran and dbt, we've built the reusable, trusted data assets that are central to how we scale AI and drive innovation.” The combined innovations are now available and will be showcased throughout Snowflake Summit 2026. Summit attendees can visit the Fivetran booth (booth #2313) or the dbt Labs booth (booth #2112) to connect with product experts, experience demonstrations of the combined innovations, and explore customer use cases powering trusted AI agents at scale. [Learn how](http://www.fivetran.com/blog/fivetran-dbt-an-open-agent-ready-future-for-data-teams) Fivetran + dbt Labs are building the data foundation for trusted AI agents. **About Fivetran + dbt Labs** Fivetran and dbt Labs deliver the data infrastructure layer that makes agents trustworthy — from the moment data moves, through every transformation, to the context an agent reasons from. The Fivetran platform moves, manages, and transforms data from every system a business runs on into a secure, reliable foundation engineered to evolve, with the flexibility to work across clouds, engines, and tools. With Fivetran, analytics, operations, and AI run on data you trust and control. Thousands of organizations worldwide, including OpenAI, LVMH, Pfizer, and Verizon, rely on Fivetran to turn data into a competitive advantage. Learn more at [Fivetran.com](http://fivetran.com), or follow Fivetran on [LinkedIn](http://linkedin.com/company/fivetran). Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 100,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://www.linkedin.com/company/dbtlabs/mycompany/), [X](https://x.com/dbt_labs), [Instagram](https://www.instagram.com/dbt_labs/), and [YouTube](https://www.youtube.com/c/dbt-labs). --- --- title: "What data infrastructure do agents need?" description: "Discover the data infrastructure architecture AI agents need to eliminate hallucinations, handle sub-second freshness, and query" url: "https://www.getdbt.com/blog/what-data-infrastructure-do-agents-need" date: "2026-05-28" authors: ["Joey Gault"] categories: ["Pulse"] --- # What data infrastructure do agents need? AI agents do not fail because of a weak model. Rather, they fail because of using infrastructure that was never meant for autonomous decision-making. As enterprises shift from simple chatbots to autonomous agents executing multi-step business processes, data infrastructure requires a fundamental change. Traditional AI infrastructure built for batch model training and human analytics fails in agentic workflows. It becomes a bottleneck for systems whose success depends on freshness, low latency, and semantic consistency. Production-ready [AI agents](https://www.getdbt.com/product/dbt-wizard) need an adaptive [data infrastructure](https://www.getdbt.com/blog/data-infrastructure) that combines real-time ingestion with a centralized semantic layer and standardized interfaces for task execution. In this article, we'll discuss what makes agent data infrastructure different and how to build it. We'll also look at the tools you can use to build a foundation that agents can trust. ## What makes agent data infrastructure different? Traditional AI infrastructure supports passive, chat-based retrieval, where systems tolerate high latency and rely on loosely structured documents to generate text. Autonomous agents operate as active software applications that execute multi-step workflows, invoke tools, trigger APIs, and retrieve specific records to drive real-world business actions. This shift from passive text generation to active execution introduces three key differences from standard AI infrastructure. - Freshness requirements. Agents demand sub-minute data freshness. A fraud detection agent cannot rely on hours-old data to approve live transactions. Without instant access to current status and support tickets, agents risk making incorrect decisions based on conditions that no longer exist. - Access patterns. Agents require diverse, concurrent access patterns. While humans typically run single analytical queries, agents use multimodal retrieval and combine vector searches, SQL queries, and key-value lookups in one workflow. Infrastructure must support these complex patterns instantly without performance degradation. - Semantic grounding. Agents need rigorous semantic grounding to avoid hallucinating definitions. Data must include metadata and governed definitions so agents can use exact programmatic logic without guessing relationship keys. ## Key infrastructure gaps that cause agent failure in production Deploying autonomous agents on legacy infrastructure leads to numerous failure modes that impair reasoning and cause systemic logic failures. - The context vacuum. It occurs when raw database tables lack metadata or documentation, leaving agents unable to interpret field relationships. For example, without machine-readable logic to define status codes, agents must infer meanings from generalized training data. This results in operational failures when AI assumptions conflict with enterprise realities. - Lineage blind spots. When agents lack end-to-end lineage, they cannot trace where a data field originated, which transformations it passed through, or whether an upstream pipeline change has corrupted it. This makes it impossible to evaluate whether the data they are acting on is reliable. - Metric drift. Siloed metric definitions across business units produce conflicting agent reasoning. When the finance team defines "active customer" differently from the product team, an agent pulling from both systems will generate outputs that contradict each other. Those contradictions surface as unpredictable operational outcomes that are difficult to trace back to a root cause. - Integration overhead. Building, securing, and maintaining custom API endpoints for every new agent use case slows deployment momentum. Teams spend more time on integration plumbing than on agent logic, and each custom endpoint introduces a new failure point that requires monitoring and maintenance. ## Building the five-layer agent data architecture: A step-by-step blueprint To move AI initiatives from pilot to production, [data engineering](https://www.getdbt.com/blog/what-is-data-engineering) teams must implement a five-layer data architecture optimized for machine consumption. ### Step 1: The ingestion layer (transition from batch to log-based CDC) Implement real-time data pipelines to capture changes in source data the moment they occur. This removes the processing lag of [batch ETL pipelines](https://www.getdbt.com/blog/etl-pipeline-best-practices) and prevents agents from acting on outdated warehouse snapshots. Log-based [Change Data Capture (CDC)](https://docs.getdbt.com/blog/change-data-capture) is the optimal approach for the ingestion layer. It reads database transaction logs to identify modifications such as inserts, updates, and deletes. Each change in the log is linked to an ordered log sequence number (LSN) that helps the CDC system determine the sequence of modifications. The CDC approach tracks changes within milliseconds of the database committing the transaction and streams these changes as discrete events to downstream destinations. This sub-second latency ensures agent data remains aligned with operational reality in near real-time. ### Step 2: The processing and context store layer (reshape streams for rapid retrieval) Raw CDC streams are often too technical for direct use. The processing layer applies the transformations to filter Personally Identifiable Information (PII), manage schema changes, and consolidate transactional events into business entities. Processed data then feeds low-latency environments such as Redis and vector databases. This multi-store strategy lets agents select the retrieval method that best matches their current task. For example, if an agent needs to recall a session history, it queries a fast NoSQL key-value store. Similarly, if it needs to match a natural language query to a technical manual, it performs a similarity search in the vector database. The processing layer provides high-velocity data through real-time data pipelines. This ensures information is optimized for live model queries. It also reduces delays associated with batch-style reporting. ### Step 3: The semantic layer (centralize meaning in a machine-readable store) Most traditional [data pipelines](https://www.getdbt.com/blog/data-pipelines) stop after loading data into a warehouse. While they govern the [data transformation](https://www.getdbt.com/blog/data-transformation), they leave the data retrieval logic ungoverned. Without [governance](https://www.getdbt.com/product/governance), AI agents may misinterpret the necessary join logic and aggregation methods. The [semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) mitigates the issue by centralizing meaning in a machine-readable format. It standardizes logical concepts so agents do not misinterpret data grains or confuse conflicting column names. Semantic layer uses entities, dimensions, and measures to create a semantic graph. When an agent requests a business metric, the semantic layer automatically references the graph to find the best path between tables and generates optimized SQL for accurate results. This method avoids fan-out queries, chasm joins, and agents from retrieving inaccurate results. ### Step 4: The validation layer (instrument pipelines with automated quality checks) Since AI agents process data as probabilistic engines, poor data leads to confident but corrupted outputs. The validation layer acts as a circuit breaker, preventing anomalous data from reaching inference engines and causing hallucinations. The validation layer runs testing queries during the data build phase. These checks use [SQL SELECT statements](https://docs.getdbt.com/sql-reference/select) to verify schema integrity, check for null constraints, and monitor distribution shifts. If the test returns zero failing rows, the data passes validation. But if failures are detected, the pipeline halts and quarantines data to protect the agent's context window. ### Step 5: The interface layer (standardize access control via open protocols) Connecting AI systems to enterprise data has long suffered from an integration explosion known as the N × M problem. If an enterprise uses four different AI models and needs to connect them to four different internal data sources, teams must build and maintain sixteen separate integrations. Every custom integration requires unique OAuth flows, distinct message parsing logic, isolated error handling routines, and custom rate limiting. The interface layer replaces this fragile, hard-coded custom API integration with open connectivity frameworks like the [Model Context Protocol (MCP)](https://www.getdbt.com/blog/mcp). MCP acts as a universal transport protocol for AI and enables teams to build a single server for a specific enterprise tool or data source. Any MCP-compliant AI model connects to that server instantly using standard JSON-RPC 2.0 messages. This standard allows autonomous agents to discover, reason about, and query corporate data assets through a unified interface. It ensures that the organization enforces strict governance and access policies before transmitting any data to the agent. ## How dbt helps you build a data infrastructure that agents can trust Implementing the five layers is difficult through custom tools. dbt provides that foundation out of the box. ### MCP Server The [dbt MCP Server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) uses the Model Context Protocol to expose [dbt models](https://docs.getdbt.com/docs/build/models), [metrics](https://docs.getdbt.com/docs/build/build-metrics-intro), and semantic definitions to AI agents without custom APIs. This open standard establishes a secure, restricted interface for LLMs to access governed corporate data through standardized host, client, and server roles. Data teams can deploy the [dbt MCP Server](https://github.com/dbt-labs/dbt-mcp) using two architectures: - Local Server: It runs on developer machines via uvx dbt-mcp and integrates with [dbt Core](https://docs.getdbt.com/blog/dbt-core-v2-is-here?version=2.0&name=Fusion) for local testing and documentation. - Remote Server: The remote server connects to [dbt platform](https://www.getdbt.com/product/dbt) via HTTP for high-scale production use, enabling multi-user metadata discovery and lineage tracing. ### Semantic Layer The[ dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) centralizes business metric definitions using [MetricFlow](https://docs.getdbt.com/docs/build/about-metricflow). Every agent queries the same governed version of business logic regardless of which tool or workflow it operates from. It removes the metric drift and context vacuum gaps that cause the most persistent agent failures in production. ### Lineage and metadata dbt maps projects as [Directed Acyclic Graphs (DAGs)](https://www.getdbt.com/blog/dag-use-cases-and-best-practices), providing agents with model relationships and source metadata. This metadata transfers to the agent through the [Discovery API](https://docs.getdbt.com/docs/dbt-apis/discovery-api?version=2.0&name=Fusion), giving the agent a logical map to discover datasets, trace column-level lineage, and audit the origin of its inputs. ### Fast, low-cost development with the dbt Fusion engine The[ dbt Fusion engine](https://docs.getdbt.com/docs/fusion) validates SQL models against your data definitions locally as you write them. It parses projects up to 30x faster than dbt Core and cuts warehouse compute costs by up to 30%, so data teams can iterate quickly and ship agent-ready data without the build/break/fix cycles that slow production deployments. ### Built-in data quality dbt implements [data quality checks](https://docs.getdbt.com/docs/build/data-tests) at transformation time, so that agents never act on corrupted records. It supports singular data tests, like custom one-off SQL queries that check highly specific business logic, and generic data tests, which function as parameterized, reusable macros applied across multiple columns via YAML configurations. ## Conclusion Traditional batch data architectures cannot support autonomous agents due to a lack of data freshness, semantic grounding, and standard interfaces. Using raw data without metadata causes agent model hallucinations and broken pipelines. An adaptive data infrastructure architecture offers a reliable, machine-readable environment for agents. It centralizes semantics, automates validation, and standardizes interfaces, giving agents the reliable foundation they need to succeed in production environments. The dbt platform provides a governed, semantically consistent data foundation that agents need to operate reliably from day one.[Get started with dbt](https://www.getdbt.com/) today. --- --- title: "Are you ready for the dbt Fusion engine?" description: "A practical look at how Brooklyn Data’s Fusion Readiness Assessment helps teams plan a migration with confidence." url: "https://www.getdbt.com/blog/are-you-ready-for-the-dbt-fusion-engine" date: "2026-05-20" authors: ["Michael Carlone"] categories: ["Insights"] --- # Are you ready for the dbt Fusion engine? _This guest post comes from Michael Carlone at Brooklyn Data._ The dbt Fusion engine raises the ceiling on what data teams can build, and how fast they can build it. Fusion delivers faster, more responsive dbt development. It catches SQL errors in your IDE before anything hits the warehouse, and gives AI coding tools the project context they need to generate accurate code. For most organizations, the question is not whether Fusion is valuable. It is whether their current foundation can support the move successfully. That is why Brooklyn Data built a **Fusion Readiness Assessment**. It helps teams evaluate what has to be true across their project, their delivery practices, and their team structure to migrate with confidence instead of guesswork. ## Why readiness matters now Without the dbt Fusion engine, teams working with dbt often rely on warehouse compute just to catch issues that are relatively small: syntax errors, broken references, and type mismatches. That slows development down and adds cost to work that is often just part of normal iteration. Fusion improves that loop by giving teams faster feedback as they work. The result is a development experience that feels quicker, more responsive, and better suited to the way modern analytics teams build, and is designed for an agentic development experience. That upside is exactly why readiness matters. A move to Fusion does not only test the engine. It also tests the broader system around it: project design, development and deployment habits, governance, ownership, and technical depth. ## What “Fusion ready”‌ means Brooklyn Data defines Fusion readiness in a simple way: you are Fusion ready when you have confidence that your data foundation can support the move smoothly, securely, and at scale. That confidence should not come from instinct alone. It should come from a clear view of how your organization works today and where migration friction is most likely to appear. That is also why readiness is not a universal checklist. Every data team operates with a different level of maturity, a different tolerance for change, and a different mix of technical and organizational complexity. A good assessment should reflect these realities. ## How the assessment works Our assessment looks at readiness across three pillars: **Project, Process, and People**. We use these three because, in practice, they shape the outcome of nearly every dbt migration. Looking at all three pillars helps teams avoid blind spots. Migration risk rarely lives in one place, and focusing only on the dbt project itself can cause teams to miss process gaps, unclear ownership, or concentrated expertise that could slow the migration once technical work begins. Even when the technical foundation is ready, the transition can still become harder to manage if deployment practices are inconsistent, governance expectations are unclear, or key knowledge sits with only a few people. Assessing Project, Process, and People together gives teams a more complete view of what needs attention before they move forward. ### Project Within **Project**, we look at the technical patterns that affect migration complexity. That includes how standardized the project is, how much custom logic is in play, how heavily the team relies on advanced features, and whether the broader foundation reflects dbt best practices. ### Process Within **Process**, we look at governance, delivery, and repeatability. Can the team develop, review, deploy, and govern work in a way that reduces risk? Are the workflows consistent enough to support change without introducing avoidable instability? ### People Within **People**, we look at organizational clarity and depth. Does the team know who owns what? Do they have the hands-on knowledge required to support a migration? Is the organization positioned to learn, adapt, and handle issues without over-relying on a single person? Each pillar is assessed, quantified, and then rolled into an overall Fusion Readiness score. The goal is to make the output easy to understand and useful in conversation, not to create a false sense of precision. In practice, the assessment is meant to answer a short list of planning questions: - How ready are we today? - Where is migration friction most likely to show up? - What should we focus on first? - What should improve before we reassess? ## Why the pillars need to be read together The assessment looks at Project, Process, and People separately, but the real value comes from reading them together. A dbt project may be technically mature, but if delivery practices are inconsistent, migration can still become harder to manage. A team may have established governance and release patterns, but if ownership is unclear or hands-on expertise is uneven, execution can still slow down. And a capable team should not have to compensate for project issues that could be surfaced and addressed earlier in the process. That is why we do not treat any one pillar as the answer. Fusion readiness is shaped by how the technical foundation, the operating model, and the team support one another. Looking at them together gives a more realistic view of where migration will be smooth, where friction is likely to appear, and where focused improvement will have the biggest impact. ## What the output helps teams do The value of the assessment is not just the score. It is the interpretation behind it. When the results come back, teams can see where readiness is more established, where there are gaps to address, and what should be prioritized before moving forward. In some cases, that may point to project cleanup. In others, it may highlight process refinement or the need for clearer ownership and broader hands-on expertise. That makes the output useful at multiple levels. Practitioners get a clearer view of the technical and operational work that will reduce migration friction. Leaders get a shared language for planning, prioritization, and risk management. Instead of relying on instinct or general enthusiasm, teams can make decisions based on a more complete view of their starting point. A good readiness assessment does not just tell you where you stand. It helps you decide what to do next. Fusion is the destination. Readiness is the roadmap. Fusion represents a meaningful step forward for dbt teams. The speed improvements are real. The developer experience is materially better. The long-term upside is clear. But the teams that will benefit most are not the ones that move first. They are the ones that understand what they are moving from, what needs attention before the transition, and how to make that move with confidence. That is the role of readiness. Our Fusion Readiness Assessment helps teams take that first look inward, quantify what matters across Project, Process, and People, and turn that picture into a more practical migration plan. If your team is thinking seriously about Fusion, the smartest first step is not to assume readiness. It is to assess it. ## Want to know your Fusion Readiness score? We can walk through the assessment with your team, quantify readiness across the three pillars, and identify the next steps that will make a migration smoother and lower risk. [Measure your readiness here](https://www.brooklyndata.co/campaigns/dbt-readiness-assessment) --- --- title: "Get dbt certified. Stay certified. Stay ahead." description: "Get dbt certified -- and stay that way. Here's why certification matters for your career and how to earn it." url: "https://www.getdbt.com/blog/get-dbt-certified-stay-certified-stay-ahead" date: "2026-05-20" authors: ["Laurent Goldsztejn"] categories: ["Learn"] --- # Get dbt certified. Stay certified. Stay ahead. dbt keeps moving. The practitioners who build with it should move with it. Analytics engineering has matured quickly. Employers expect more, projects are more complex, and the bar for what good looks like keeps rising. If you work with dbt, getting certified isn't just a nice-to-have. It's how you prove you're operating at the standard the industry now demands. ## Why get certified? Because credibility matters, and proof beats claims every time. The [**dbt Labs certification program**](https://www.getdbt.com/dbt-certification) offers two exams designed for the practitioners actually building data pipelines, writing models, and owning data quality in production. Whether you're an analytics engineer, a data analyst moving into engineering, or a dbt platform power user, there's a path built for where you are. Passing the exam tells employers, clients, and collaborators something a resume line can't: you've been tested against a real standard, and you passed. It helps you: - Stand out as an expert in a crowded field of dbt practitioners - Build trust faster with new teams and clients - Move into senior roles with a credential that backs you up - Show your employer that their investment in you is paying off as you continue learning ## Two exams. One standard. The [**dbt Analytics Engineering Certification**](https://www.getdbt.com/certifications/analytics-engineer-certification-exam-version-1-11) tests your ability to apply dbt's core framework—modeling, testing, documentation, and deployment—the way it's actually done in production environments. The [**dbt Architect Certification**](https://www.getdbt.com/certifications/dbt-architect-certification-exam) tests your ability to design secure, scalable dbt implementations, with a focus on environment orchestration, role-based access control, integrations with other tools, and collaborative development workflows aligned with best practices. Together, they cover the full range of what it means to work with dbt at a high level. By getting certified, you’ll join 3,100+ other dbt developers and architects who have earned the credential. ## Get certified on dbt Because relevance matters as much as experience. dbt evolves. Best practices shift. An active certification tells the world you're keeping pace, and not coasting on knowledge from two years ago. If you let your certification lapse, it doesn't erase what you know. But it weakens the signal at exactly the moment someone is deciding whether to trust you with their data stack. Stay certified to show you're current, not just experienced. Committed, not complacent. ## This is your moment. This isn't starting over. With the [**dbt Analytics Engineering Exam**](https://www.getdbt.com/certifications/analytics-engineer-certification-exam-version-1-11) updated to align with [dbt Core 1.11](https://www.getdbt.com/blog/dbt-core-v1-11-is-ga), it's your chance to close any gaps, get hands-on with what's changed, and prove your skills against the standard that matters today. To sharpen your preparation, we’re including 10 sample questions on the web page and in the study guide so you know what to expect before you sit for the exam. And when you pass the exam again to regain your credential? It'll mean even more. ## Keep your edge Certification isn't a one-and-done achievement. It's a standard you maintain, and a signal to everyone you work with that you take this craft seriously. So whether you're earning it for the first time or earning it again: Stay certified. Stay relevant. Stay ahead. [Stand apart with dbt Certification.](https://www.getdbt.com/dbt-certification) --- --- title: "AI-ready data in practice: What dbt Semantic Layer and dbt's MCP server and agent skills do for your team" description: "Clean data is just the start. See how dbt's semantic layer, MCP, and agent skills give AI the business context it needs." url: "https://www.getdbt.com/blog/ai-ready-data-in-practice-what-dbt-semantic-layer-and-dbt-s-mcp-server-and-agent-skills-do-for" date: "2026-05-19" authors: ["Stephen Thibeault"] categories: ["Insights"] --- # AI-ready data in practice: What dbt Semantic Layer and dbt's MCP server and agent skills do for your team When it comes to getting their data AI-ready, many organizations start with cleaning and structuring their data and then simply stop. This is an important first step, but it’s not the last step, because AI-ready data relies heavily on context: the layer of meaning that explains what your data‌ actually represents. You need to gather as much information as you‌ can about that data: Where are data points coming from? Which team defines the metric? Which team owns inputting this data into a system? Without answers to questions like these, even clean, well-structured data can lead AI astray. One way to think about AI is as a great teammate that knows SQL and analytics really, really well but knows zero about your organization. An agent doesn't know the different acronyms used in your industry, for example, and it doesn’t understand your business goals. For AI to work effectively and efficiently, you need to give it all that important context to make the data meaningful. In practice, teams use dbt’s AI capabilities to make data meaningful to AI agents. dbt lives on top of tools like Snowflake, BigQuery, and Databricks to transform data without having to use stored procedures or other data transformation techniques, and there are three key pieces to dbt’s AI stack: the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=1.12), and [dbt agent skills](https://docs.getdbt.com/blog/dbt-agent-skills?version=1.12). Here’s what they are, how they work together, and how to use them to ensure high-quality, AI-ready data. ## The semantic data layer is your lens The semantic data layer provides all of the context that the AI will need to understand your data: the structure of the data, how you work with the data, and what exists in the data. I think of it like this: I have very bad eyesight. When I take my glasses off, I can still see things, but they are far from in focus. There will be some things that I miss and other things that are incomplete in my vision because I can't fully see everything. When I put my glasses on, I'm able to see clearly and completely. This is essentially what a semantic layer does for your data. A generic semantic layer is like buying plain, off-the-rack reading glasses. It makes things somewhat clearer; you will get answers some, but not all, of the time, and you’re not getting the most detailed vision possible. A governed, [dbt-backed semantic layer](https://www.getdbt.com/blog/semantic-layer-introduction) gives you prescription lenses that are custom-focused for your business's vision, signed off by someone trusted, and updated through scheduled exams as your vision (your data, your definitions, your business) change. AI wearing drugstore readers might see something somewhat clearly, but it'll squint and need to occasionally guess. AI wearing your prescription sees exactly what your business means by "revenue," "active customer," or "churn" and keeps seeing correctly as those definitions evolve. So when we talk about gathering context around data, most of that context is typically handled within the semantic layer. This is especially true when it comes to what certain columns mean, what certain metrics are, and how different values or properties are to be calculated. ### You don't need a perfect semantic layer to start You can get a lot of use out of dbt’s AI tooling even without a semantic data layer in place. The semantic layer is mainly used for conversational AI, letting agents query your actual data and return reliable AI outputs. But if you want to use dbt's AI tooling for development workflows, you don't need it. There are still things that you can do with dbt's AI tools outside of it, like diagnosing job failures, finding column-level lineage, and other things that really speed up your workflow. Don't let not having your data fully cleaned up, or not yet having your data fully defined in the semantic layer, be what stops you from using dbt’s AI tools. You can absolutely start using them now, and you can even use some of them to help build your semantic layer as you go. ## Three pieces of the dbt AI stack Terms like "agent skills" and "MCP server" can be‌ intimidating when you first hear them. Let's demystify these. **MCP server: the tools.** An MCP server is a set of tools like API calls that can be used to communicate with applications on the backend. Its function is to give the agent instructions on how to make those calls and how to use what it gets back. For example, there's a tool called **list_metrics** used to pull data from the semantic layer, and another one called **get_job_run_error** for diagnosing failures available as functions in the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=1.12). The dbt MCP server grounds those interactions in structured, dbt-native context, so agents are working from what your data actually means, not guessing from static documentation. **Agent skills: the instructions.** [dbt’s agent skills](https://docs.getdbt.com/blog/dbt-agent-skills?version=1.12) are workflow instructions that give your agent proven, opinion guidance for common dbt tasks like writing tests, debugging failures, defining metrics, handling migrations. They load on demand and only when relevant. An agent skill gives the agent a set of clear instructions needed to complete a specific task. Skills also provide the agent with rules and guardrails: never do this; here are common pitfalls you may run into; here are things that you need to look out for. **How the semantic layer, MCP, and agent skills fit together:** Each piece has a distinct role, and together they cover everything an agent needs to work effectively with your data. The semantic layer provides the context, MCP provides governed access, and agent skills provide the proven workflows agents need to query the data or to get the tools they need out of the MCP server. ## dbt’s AI tools in production to speed up data development The best way to understand how these three pieces work together is to see them in action. One of our clients, a very large technology company, used them to feed structured data into a Slack channel where dbt errors are automatically sent. They hooked dbt's MCP server, along with Claude, into that error triage channel to look at those job failures and actually diagnose them. The integration uses the **get_job_failure** function in the dbt MCP server, looks at the error, and then has the agent analyze what happened and why. By the time a developer actually gets to that error they're able to see a quick triage that was already done, along with some possible solutions. This integration is not fully set up for self-healing just yet. There are definitely controls around the AI, and it doesn't get everything right all of the time, but it's a huge time save. Instead of having to go into the dbt platform and dig through the logs to find the specific problem, you have it all laid out there by your agent. That same team is also working on a GitHub action: if somebody creates a model and doesn't include a semantic layer definition, the agent will try to create one and send it back to the developer with a note: _here's what I created, add on to it to make your semantic layer._ The goal is to encourage that hygiene of getting that context as a natural part of the workflow, rather than an afterthought. And, notably, both of these are use cases that don't require a semantic layer at all. ## Where to start: pilot small and smart If you're ready to include AI in your data pipelines, the most important advice I can give is to do it in steps. Really hone in on one business unit that is willing to work with you on a pilot program for AI readiness, and focus on gathering semantics around the data for that small subset. (Pilots within the data team itself, like the error triage example above, are a great place to start. They can be very useful, and they don't require a well-crafted semantic layer to work effectively. So there's no reason to wait!) Gathering that semantic information, though, will really allow you to get your feet under you when it comes to building a semantic layer, and it will allow you to iterate very quickly. When you collaborate with one team in a pilot project, you're able to break things and learn from your mistakes before bringing it out to more business units. So: start small, really focus in on what you're able to do (and what you reasonably _can_ do), and then apply what you learned. Then you can use the momentum you gain by providing something great to that particular team or business unit to expand the semantic layer to more teams across your org. ## Why semantic standards matter: Open Semantic Interchange Once you’re ready to build out your semantic layer it’s important to understand that, right now, basically every data tool implements semantic definitions in its own proprietary format.. Power BI has one, Omni has one, Databricks has one, Snowflake has one, and of course [we have one](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl?version=1.12). That fragmentation creates a portability problem: if your semantic layer definitions (metric names, calculations, business logic) are expressed in a format that’s specific to one tool, you can't move them to another tool without rebuilding everything from scratch. So, for example, if you define "monthly recurring revenue" in dbt's Semantic Layer and then want to also expose that definition in Power BI or Snowflake, you'd have to redefine it natively in each system. Besides redundancy it also creates inconsistency risk and a lot of maintenance overhead. This is why dbt, along with Snowflake, Databricks, and a large number of other major organizations in data, have joined an initiative called the [Open Semantic Interchange](https://open-semantic-interchange.org/). The v0.1 a vendor-neutral spec is already live and open source. It’s an industry-wide specification that standardizes how we exchange semantic metadata across analytics, AI and BI platforms. The OSI spec serves as a common language for metrics, dimensions, and relationships, so metrics can be interpreted consistently across tools (e.g., Snowflake, Tableau, dbt) while minimizing vendor lock-in. The dbt Semantic Layer complements the spec by making those definitions operational: you define and govern your metrics in the dbt Semantic Layer using MetricFlow, and OSI provides the interchange format to move those definitions across other tools like Snowflake and Tableau. Author once, use everywhere. Think of it‌ like the same reason we needed MCP in the first place: when there's no common standard, every tool reinvents the wheel and nothing moves cleanly between systems. A shared standard changes that. --- --- title: "What's shipped in dbt — May 2026" description: "A roundup of everything we've shipped since January—across agents, Fusion, security, developer experience, dbt Core, and more." url: "https://www.getdbt.com/blog/what-s-shipped-in-dbt-may-2026" date: "2026-05-19" authors: ["Corinne Hallander"] categories: ["Product"] --- # What's shipped in dbt — May 2026 It's been a big few months of shipping at dbt. We've got a lot to cover — from the dbt Developer Agent going into preview, to making the upgrade to the dbt Fusion engine self-serve, to new ways to lock down your account security, to quality-of-life improvements for practitioners who live in the IDE. Here's everything that's landed since January. ## AI that works with your data, not around it ### dbt gets an AI-native developer: the dbt Developer Agent (Preview) General-purpose coding agents are now everywhere, ready to help anyone code. But the question we kept hearing from teams this year was some version of: can we get an agent that actually works like an analytics engineer? One that‌ understands my whole dbt project? One that can read the graphs, knows the lineage, validates before it touches anything, and helps me build dbt models without breaking anything? This is why we’ve built the dbt Developer Agent, which is now available in Preview for dbt platform customers with dbt Copilot enabled. Simply describe the change you want to make — rename a model, add a metric, migrate a stored procedure, fix a failing build — and the agent reads your graph, understands what's upstream and downstream, and drafts the edits across every file that needs to move. SQL, YAML configs, tests, documentation: coordinated changes in one pass. That means less time context-switching between files, fewer broken builds, and data work that ships faster. → [Read our full announcement blog to learn more](https://www.getdbt.com/blog/the-dbt-developer-agent-is-now-in-preview) ### dbt Agent Skills - GA Earlier this year we released [dbt Agent Skills](https://github.com/dbt-labs/dbt-agent-skills) — an open-source repository of best practices that teach generalist coding agents how to think like an analytics engineer that actually understands how to work with dbt projects. Skills are structured knowledge files that agents load on demand. They encode things like: when to preview data before writing tests, how to structure a semantic model, how to debug a job failure without chasing the wrong root cause. Check out our growing repository of skills by clicking below: → [dbt Agent Skills on GitHub](https://github.com/dbt-labs/dbt-agent-skills) ### Securely connect dbt to your favorite AI tools (Beta) The dbt MCP server now supports OAuth, so you can now connect OAuth-enabled AI tools — Claude, ChatGPT, Glean, and others — to dbt using your existing dbt login. No token management, no configuration hand-off to an admin. Your identity, properly permissioned and secure, in a few clicks. [→ OAuth integrations docs](https://docs.getdbt.com/docs/platform/manage-access/connect-apps-oauth) ### Remote MCP Server: Admin API support + product docs tools Two new sets of tools landed in the dbt Remote MCP Server. First, the MCP server now supports Admin API calls — which means AI assistants (Claude, Cursor, etc.) can help troubleshoot job errors directly, not just write queries. Second, the MCP server now includes search_product_docs and get_product_doc_pages tools that pull from docs.getdbt.com in real time, so you get answers grounded in the actual docs rather than training data. → [dbt MCP repo](https://github.com/dbt-labs/dbt-mcp) ### Bring your own Anthropic key dbt Copilot now supports BYOK (bring your own key) for Anthropic, so teams can power their AI workflows in the dbt platform using their own Anthropic API key — with the usage, cost, and data handling that comes with it. BYOK is also available for OpenAI and Azure OpenAI, giving teams flexibility to build with the model provider that fits their security, compliance, and cost requirements. → [Read the docs to learn more](https://docs.getdbt.com/docs/platform/enable-dbt-copilot#configure-your-ai-provider) ## Getting to Fusion just got a lot easier The big headline on the Fusion side this cycle is that adoption is now self-serve in dbt platform. By accelerating your upgrade to Fusion, you can take advantage of 30x faster parsing time, richer metadata for AI, realtime feedback on SQL as you type, and more. But upgrading your projects manually one-by-one could take hours or days…why not let dbt do the hard parts for you? ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/713a573e9671b3ed3b7351cf89a66b8499c12b92-1478x926.jpg) ### Upgrade to Fusion project by project If you're a dbt platform customer, you can now see which of your projects are eligible for Fusion and move them one at a time, directly from the platform UI. Pick a project, follow the prompts – no ticket, no wait, no overhead. ### Fusion migration skill (Beta) Upgrading to Fusion shouldn't mean fixing conformance errors manually. The Fusion migration skill in the dbt Developer Agent brings an automated approach to getting your projects Fusion-ready, faster: - It classifies every conformance failure, - Applies only validated high-confidence fixes automatically, and - Walks you through medium-confidence changes with clear diffs and your approval. Blocked issues– those caused by Fusion bugs or framework limitations– are surfaced immediately, with context and a path forward. No wasted effort chasing unfixable errors. The skill re-validates after every fix to handle cascading errors correctly and ends every session with a transparent report. This gives you faster triage, safer fixes, and trust in your upgrade. **How to get started:** 1. **In dbt Studio: **Find a job or project that’s ineligible for Fusion. Attempt the Fusion run so you can see the build conformance errors. **Studio will surface a new entry point directly in that conformance error experience** so you don’t have to dig through error logs. From there, launch the conformance skill and enjoy! 2. **Via VS Code: **The dbt VS Code extension now makes Fusion setup and upgrade significantly easier. When you're ready to upgrade your project, you can run the CLI onboarding flow in the terminal or let an AI agent handle it via the dbt Agent Developer or Cursor, no command line required. Start your seamless upgrade to the dbt Fusion engine: → [Learn more](https://docs.getdbt.com/guides/upgrade-to-fusion?step=3#step-1-start-the-upgrade-assistant) ### More from Fusion this cycle: Beyond easier adoption, we've invested in making the engine faster and more capable. - **UDF-aware deferral.** When you run with --defer and --state, dbt now resolves function() calls from the state manifest — so models that depend on UDFs don't require you to rebuild those functions in your current target first. - **Python UDFs** are now supported on Snowflake and BigQuery in the Fusion engine CLI. - **DuckDB support** **(Beta)**. Run local dbt projects without a warehouse account. Useful for testing, exploration, and CI scenarios where warehouse costs matter. - **Apache Spark 3.0 (Beta)**. Fusion engine CLI support for Spark means faster compilation and execution for Spark-based dbt projects – no Python runtime, no subprocess overhead. For dbt platform customers: - **dbt compare** **from local dev to CI**. You can now compare changes at every stage of your workflow. In local development, the dbt VS Code extension previews how your edits affect your data (added/removed rows, join verification) before you open a PR. Then at the CI stage, dbt compare runs in orchestration on Fusion, giving you model-level diffs as part of your pipeline gate automatically. - [**Fusion release tracks**](https://docs.getdbt.com/docs/dbt-versions/cloud-release-tracks?version=2.0#fusion-release-tracks) give you control over your update cadence: Nightly, Stable, Extended, and Fallback. Choose the release track that matches your team’s stability requirements, risk tolerance and change management processes. - **New projects default to Fusion Stable**. New environments in Developer, Starter, and Enterprise accounts now provision on the “Fusion Stable” release track by default – for any supported adapter (Snowflake, Redshift, BigQuery, Databricks). Want to fast-track your migration to Fusion? Use our quickstart guide. [–> Quickstart guide for Fusion](https://docs.getdbt.com/guides/fusion?step=1) ## For dbt builders: Developer experience improvements This cycle we focused on the things practitioners have been asking for: faster navigation in the IDE, more context at a glance, broader warehouse support for query history, and a meaningfully simplified semantic layer spec. ### Studio IDE: search, replace, and command palette The Studio IDE now has search and replace across your project, a command palette, and the ability to jump to symbols and run IDE configuration commands. These capabilities have been long-requested, and now they're here. ### Studio IDE: Better status bar The status bar now surfaces deferral settings, dbt version, and project status with quicker access to change them. ### Model query history: Databricks and Redshift — Bet**a** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2f97a4089ef8dbe3d16d7510f05d251cb2bed244-2048x1117.jpg) Model query history now supports Databricks and Redshift in addition to Snowflake and BigQuery. If you're on either of those warehouses and want to understand query patterns at the model level, this is now available in beta. →[Read the docs](https://docs.getdbt.com/docs/explore/model-query-history) ### New semantic layer YAML spec ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/0391c15f8ba49aa763bf214fe2e62de15fc028ac-1698x1276.jpg) The new [semantic layer YAML specification](https://docs.getdbt.com/blog/modernizing-the-semantic-layer-spec?version=2.0) introduces several key changes: semantic models are now embedded within model YAML entries (no more managing entries across multiple files), measures are now simple metrics, and frequently-used options are promoted to top-level keys. This is a meaningful spec simplification making it easier for anyone maintaining a semantic layer, and a lower barrier to adoption for those who haven't yet. The new specification is live in dbt Core v1.12 and on the dbt platform “Latest” release track. →[ Migrate to the latest YAML spec](https://docs.getdbt.com/docs/build/latest-metrics-spec?version=2.0) ## ## Access to dbt that’s secure, governed and self-serve We shipped several updates this cycle to make security configuration simpler — and in most cases, self-serve. ### Global login — GA ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/baf87657a055672e2bb41584ad2f6d2faa415aed-2048x1183.jpg) ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2fd33a5b1272ef7ac83d2653a6d3110eda2b2cdc-1910x1080.jpg) There's now a universal login URL that shows all the accounts you have access to across regions and tenancies, in one place. This is available now for multi-tenant accounts with an account-specific domain; single-tenant support is coming soon. →[ Log in to dbt platform](https://login.dbt.com/) ### Self-serve private endpoints — Beta You can now configure Snowflake PrivateLink endpoints directly in the dbt platform without filing a support ticket. Go to **Account settings → Integrations → Private endpoints** to request and manage Snowflake PrivateLink endpoints on AWS. If establishing secure connectivity for your dbt setup has been a multi-week support ticket process, that changes now. →[Read the docs](https://docs.getdbt.com/docs/platform/secure/private-connectivity/aws/aws-snowflake) ### Connection profiles — GA ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/2d593ec401c5b3c18b3e50281e4cd60a0ed92db4-1428x1482.jpg) Profiles let you define and manage connections, credentials, and attributes for deployment environments at the project level. dbt automatically creates profiles for your existing projects and environments, so there's nothing to migrate. Useful for teams that want more structured control over how credentials and connections are organized across environments. → [About profiles](https://docs.getdbt.com/docs/platform/about-profiles) ### Account-level Slack and Microsoft Teams notifications — GA ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/aee48940b197cb070aaef24e21ebabd7ed477232-1872x542.jpg) Job notifications can now be sent to Slack and Teams channels configured at the account level, not just per-job. This makes it easier to set up centralized alerting without touching every job's configuration. Both Slack and Teams notifications are now generally available. → [Slack notifications](https://docs.getdbt.com/docs/deploy/job-notifications#slack-notifications-account) · [Teams notifications](https://docs.getdbt.com/docs/deploy/job-notifications#microsoft-teams-notifications) ## dbt Core v1.12 is here in Beta The dbt language is continuing to evolve, and dbt Core v1.12 reflects that momentum. The beta release includes contributions from across the community. ### What's in v1.12: - New on_error config to control whether downstream models run when an upstream model fails. Set on_error: continue on a model to allow downstream nodes to still attempt to execute even when it errors. - Define project variables in root-level vars.yml to reference them within dbt_project.yml or to keep dbt_project.yml slim. - New selector method (selector:my_selector) to reference a named selector from selectors.yml inside --select or --exclude to combine with other selectors, graph operators, and set operators. - Support for the new semantic layer spec simplifies how you define metrics and dimensions by embedding semantic annotations directly alongside each model. - Expansions of user-defined functions (UDFs) - Use public third-party PyPI packages in your Python UDFs with the new packages config. - Write UDF logic in javascript. - Overloaded UDFs - define multiple functions with the same name but different argument signatures. - Execute ad hoc database statements (no macro needed) with dbt run-operation --sql - Improvements to exception handling so error messages are clearer and stack traces are easier to interpret. - and more coming soon! → [Learn more in the v1.12 upgrade guide](https://docs.getdbt.com/docs/dbt-versions/core-upgrade/upgrading-to-v1.12?version=2.0) ## What’s next There's always more coming. Stay tuned on our blog for the latest announcements. In the meantime, the features above are live. If you have questions, _find us in [#product-updates](https://www.getdbt.com/community/join-the-community/) in the dbt Community Slack._ Or [contact us](https://www.getdbt.com/contact) to see what dbt can do for your data team. ### See us in San Francisco this June We’re at [Snowflake Summit June 1-4 (Booth #2112)](https://www.getdbt.com/events/snowflake-summit-2026) and [Databricks Data+AI Summit June 15-18 (Booth #430)](https://www.getdbt.com/events/databricks-summit-2026). We'll have live demos, the team on site, and a lot to show you. --- --- title: "AI-assisted analytics engineering: Docusign’s framework for scaling dbt unit testing" description: "How Docusign reduced dbt unit test authoring from 5 hours to 30 minutes using a structured AI-assisted framework." url: "https://www.getdbt.com/blog/ai-assisted-analytics-engineering-docusign-s-framework-for-scaling-dbt-unit-testing" date: "2026-05-18" authors: ["Sundar Subramanyam"] categories: ["Insights"] --- # AI-assisted analytics engineering: Docusign’s framework for scaling dbt unit testing _This guest post comes from Sundar Subramanyam, Lead Data Engineer at Docusign._ At Docusign, we support millions of customers worldwide in managing critical agreement workflows. As our analytics platform scaled to support new product launches and features, ensuring data quality before production data existed became a key challenge. Traditional dbt data tests (e.g., `not_null, unique`) rely on existing datasets. However, for new features and evolving pipelines, we needed **dbt unit tests**—tests that validate logic by mocking input data and asserting expected outputs. While powerful in theory, unit testing in dbt introduced a practical problem - **the effort required to manually author tests did not scale with the complexity of our models.** To address this, we explored a focused question: **_Can AI systematically reduce the friction of dbt unit testing?_** This led to the development of a structured approach using **GitHub Copilot (GPT-4+)** that significantly improved both **testing efficiency and adoption**. We used our own AI tooling to speed up unit test drafting by ~90%, while dbt provided the structure to govern and enforce those tests in CI, making data quality reliable even before production data existed. ### **The unit testing bottleneck** Unit testing in analytics engineering is critical for validating: - Complex `CASE` logic - Join conditions and fan-out scenarios - Filtering rules and edge cases However, the real challenge lies in the setup: For each model, engineers must: - Analyze SQL logic - Create mock input datasets - Manually compute expected outputs - Debug YAML syntax In practice, this process took **up to 5 hours per complex model**, making comprehensive testing difficult to prioritize. ## **Introducing the AI-assisted dbt unit testing framework** ![Image](https://cdn.sanity.io/images/wl0ndo6t/main/4568838730aa4732f55c55c3505344144c4c2211-510x255.png) To address this, I developed the **AI-Assisted dbt Unit Testing Framework** — a structured, human-in-the-loop methodology that leverages generative AI to automate the creation of dbt unit tests. Rather than treating AI as a replacement for engineers, this framework positions AI as an **accelerator for repetitive tasks**, while preserving human validation for correctness. ## **Framework workflow** The framework follows a repeatable, multi-step process: ### **1. Model input** Engineers provide a dbt model (e.g., `dim_customer`) containing SQL transformations. ### **2. AI interpretation** A custom AI workflow parses: - Column-level transformations - Joins and filters - Logical branches (e.g., CASE conditions) ### **3. Logic summarization** The system generates a structured understanding of: - Source tables and references - Output columns - Transformation rules ### **4. Human validation** Engineers review and confirm the interpretation before proceeding, ensuring correctness and trust. ### **5. Test case generation** The framework generates: - Positive test cases - Negative test cases - Edge-case scenarios The AI focuses heavily on generating **synthetic mock data**, including: - Null handling - Boundary conditions - Join anomalies - Temporal edge cases ### **6. YAML output** The system produces a valid dbt unit test file (`*_unit_test.yml`) with: - Mock input datasets - Expected outputs - dbt-compliant structure ### **7. Iterative refinement** Engineers refine the generated tests and commit them into the dbt CI/CD pipeline. ### ### **Prompt pattern behind the framework** The core of this workflow was a structured prompt pattern rather than a one-off AI request. The prompt guided the AI through a repeatable sequence: - Interpret the dbt model logic. - Identify source references and output columns. - Summarize the logic and ask the engineer to validate the understanding. - Generate positive and negative unit test scenarios. - Create mock input data and expected outputs. - Ensure the `expect` section matches the model output columns. - Output the result as a dbt-compliant `_unit_test.yml` file. - Allow the engineer to refine the test cases through feedback. This structure helped make the workflow repeatable and reviewable, while keeping the engineer responsible for validating business logic and expected outcomes. ### **From SQL to test case** The real power is seeing how the AI handles mocking data. **Model SQL (Snippet):** SQL `CASE` `WHEN subscription_status = 'Active' AND renewal_date < current_date THEN 'Overdue'` `WHEN subscription_status = 'Active' THEN 'Current'` `ELSE 'Inactive'` `END as derived_status` AI-generated unit test: The framework generates test scenarios such as: YAML `unit_tests:` `- name: test_derived_status_logic` `model: dim_subscription` `given:` `- input: ref('stg_salesforce')` `rows:` `- {subscription_status: 'Active', renewal_date: '2023-01-01'} # Scenario 1: Overdue` `- {subscription_status: 'Active', renewal_date: '2025-01-01'} # Scenario 2: Current` `- {subscription_status: 'Pending', renewal_date: '2025-01-01'} # Scenario 3: Inactive` `expect:` `- rows:` `- {derived_status: 'Overdue'}` `- {derived_status: 'Current'}` `- {derived_status: 'Inactive'}` _The key advantage - The AI identifies logic branches and automatically generates test data to validate each scenario._ ### **Impact: 10x productivity and test coverage** The results of this small experiment were immediate and measurable: - **90% Reduction in cycle time**: Writing a comprehensive Unit test suite dropped from 5 hours to roughly 30 minutes. Engineers no longer start from a blank file - but they start with a working draft. - **Increased test coverage**: Because testing became easier, engineers tested more. We closed the gaps on edge cases that used to slip through manual review. - **Shift-left quality**: We caught complex logic bugs (mismatched joins, bad filters) locally, long before they reached the production dashboards. Implementing unit tests helped us catch at least 5–10 data defects that would otherwise have gone unnoticed. - **Scalable Trust**: Whether refactoring legacy code or building net-new models for new product and feature launches, we established a consistent baseline of quality without burning out the team. ### **What worked and what didn’t** **Where AI excelled:** - Parsing Jinja and SQL syntax to map logic branches. - Generating tedious mock data (rows of CSVs) in valid YAML format. - Identifying edge cases a human might overlook (e.g., "What if this date is null?"). **Where humans remain essential:** - Validating the _business intent_ of the logic. - Ensuring the "expected output" aligns with domain knowledge, not just code patterns. ## **Industry relevance and adoption potential ** The challenges addressed by this framework are not unique to a single organization. Many data teams struggle with: - Low adoption of unit testing - High manual effort - Inconsistent data validation practices This framework provides a **reusable and scalable approach** that can be applied across dbt projects and analytics engineering teams. ### **Looking ahead: From a win to a workflow** This initiative began as a focused experiment but has evolved into a repeatable pattern for integrating AI into analytics engineering workflows. Future directions include: - Integrating test generation into CI/CD pipelines - Generating tests from business requirements (e.g., Jira tickets) - Expanding the framework to other areas of data engineering ## **Conclusion** AI does not need to be complex to be impactful. By addressing a specific bottleneck—unit** test creation in dbt**—this framework demonstrates how targeted AI applications can deliver measurable improvements in productivity, reliability, and scalability. The broader takeaway: “Identify one friction point in your workflow—and use AI to systematically eliminate it.” --- --- title: "How Nasdaq built a governed intelligence layer with dbt and Databricks" description: "Nasdaq processes up to a trillion messages a day across 26 business lines. Here's why they use dbt and Databricks to do it." url: "https://www.getdbt.com/blog/how-nasdaq-built-a-governed-intelligence-layer-with-dbt-and-databricks" date: "2026-05-18" authors: ["Daniel Poppy"] categories: ["Insights"] --- # How Nasdaq built a governed intelligence layer with dbt and Databricks The stakes in financial services data are different from almost any other industry. In most business cases, the cost of a data error comes down to wasted time or a slightly wrong metric. In financial markets, a data break doesn't surface as an error message. It surfaces as a wrong regulatory filing, an incorrect client bill, or a risk desk making calls from numbers that were already stale. Jamie Nemeroff leads value engineering at dbt, where he quantifies what good data infrastructure is worth. In financial services, that calculation is easier to make than elsewhere because the cost of getting it wrong is so traceable. It's also what makes the story of what Nasdaq has built worth examining closely. Michael Weiss, AVP of Product at Nasdaq, leads Nasdaq Eqlipse Intelligence, an end-to-end data platform built on dbt and Databricks that Nasdaq originally developed for its own markets and now delivers to its financial market infrastructure (FMI) customers: the exchanges, clearinghouses, and central securities depositories (CSDs) that operate on Nasdaq's Eqlipse product suite. [Jamie spoke recently with Michael and Andrea DeSosa, who leads go-to-market for capital markets at Databricks, to discuss what they built, how they built it, and what it makes possible.](https://www.getdbt.com/resources/webinars/how-nasdaq-productized-a-governed-intelligence-layer-with-dbt-databricks-for-financial-market&sa=D&source=docs&ust=1778887535313379&usg=AOvVaw0hj8jcJYElXzLm3FrVUBqN) Four themes came up that Jamie tracks across every large financial services engagement he works on: scale, removing engineering bottlenecks, regulatory governance, and AI readiness. Nasdaq's story connects all four. And it's increasingly a product story: the architecture they spent a decade proving internally is now available to their customers in six months. ## Nasdaq Eqlipse Intelligence: architecture and origins Most people know Nasdaq as a market operator. A less-visible part of the company is its financial technology arm, which delivers infrastructure software to financial services firms globally, covering anti-financial crime, regulatory technology, and capital markets operations. Nasdaq's Eqlipse product suite focuses specifically on FMI customers. Michael described the intelligence platform as a data layer spanning the full lifecycle: ingestion from financial market protocols (Nasdaq-provided and third-party), transformation and mapping, validation and reconciliation, and business applications for reporting, analytics, and billing. Databricks serves as the primary computation layer. dbt plays two roles: transformation engine and [semantic layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl). On top sit three customer-facing products: InsightsHQ for visual dashboards, ReportHQ for structured report outputs like CSVs and PDFs, and RevenueHQ for managing billing and fees. The platform didn't arrive fully formed. Nasdaq has been on this path since 2012, starting with a regulatory data product for US broker-dealers. In 2014, the company moved all of its market data from on-premises warehouses to AWS. The intelligence platform was formally launched in 2021, extended to external FMI customers in 2024, and is now in active delivery to its first customers. What's notable about that timeline is the direction of proof. Nasdaq didn't build a product and then go find buyers. They built it for themselves, ran it at scale for years, and then recognized that the FMI customers they spoke with had the same problems they had already worked through. ## Scale at a trillion messages per day The volume Nasdaq manages helps explain why every architecture choice at this level carries weight. Michael described ingesting hundreds of billions of messages per day across Nasdaq's US and Nordic businesses, from thousands of data sources, across dozens of business lines. Regulatory compliance makes precision essential. US programs like the Consolidated Audit Trail (CAT) run on roughly a three-day window for getting data out. Billing has to reconcile to the exact contract. As Michael put it: "When we go to generate the bill at the end of the month, we can't have an unaccounted for or an extra contract in the bill we're sending. We need to make sure that information is right, both in terms of the input and the output." The approach Nasdaq uses to maintain consistency at this scale is what Michael described as a [dbt Mesh](https://www.getdbt.com/blog/data-mesh-architecture-explained) structure: a family of dbt projects organized from foundational models upward. The foundational layer defines contracts for raw data. Every message on an order chain looks the same, regardless of whether it comes from a Nasdaq market, a Southeast Asian exchange, or a non-Nasdaq trading platform. That consistent contract makes it possible to define metrics and KPIs once and deploy them everywhere. It also shapes how the team responds to errors. "Our stance," Michael said, "is we'd rather have the pipeline break with an error on the test and be late on sending data than send wrong data." Andrea framed scale from the Databricks side as a trust problem as much as a volume problem. "Every time a team builds its own pipeline and defines its own metrics," she said, "you've created a future audit finding, a future reconciliation break, or a future AI model trained on the wrong truth." Databricks' [Unity Catalog](https://www.databricks.com/product/unity-catalog) addresses this at the infrastructure level. The same definition of settlement finality or net position applies across all of Nasdaq's internal teams and FMI customers. That’s not because people agreed to use the same spreadsheet, but because the platform enforces it. ## Getting data to the people who need it Engineering bottlenecks are the second theme Jamie consistently work through in financial services business cases. The data exists. The business teams need it. But access runs through ticket queues, and the lag compounds. Michael described what changed at Nasdaq when they put dbt in front of non-engineering teams. Because SQL is broadly understood across the business, governed modeling tools could go directly to analysts and business users. The result was a significant acceleration in time to market for new data products and insights. The example he gave was concrete. Nasdaq's options sales team, working alongside the options business team that owned the underlying dbt models, built client-specific visuals to distribute as part of their sales process. No engineering involvement required. Nasdaq is now looking to offer the same model to its FMI customers: access to Nasdaq's foundational dbt models alongside governed tooling that lets customers' business teams build on top. "We're looking to let them take our foundational models and rebuild things from a business point of view that complement or supplement the models they need to make their business go." Andrea described why this approach compounds rather than complicates. "The biggest tax on delivery isn't talent," she said. "It's starting from scratch every time." A consistent foundation changes the unit of work: teams inherit a proven architecture and configure it to their markets. The people at FMIs who understand the business best, the clearing workflow specialists and surveillance officers, can start contributing to model development rather than waiting on engineering queues. The condition that makes this safe is governance. As Andrea said: "You can only safely put those tools in non-engineering hands when you have a governance layer enforcing the boundaries underneath." Unity Catalog handles that at the data level. dbt handles it at the model level. Together, they let more people build without fragmenting the foundation. ## Governance that holds up to regulators "Governance" can become abstract fast. In financial services, it's specific. Michael walked through two concrete examples: - Rule 17A for broker-dealer compliance in the US establishes a write-once, read-many (WORM) obligation. Firms must be able to prove that data wasn't manipulated after the fact. - SOC 2 compliance for billing means being able to show an auditor the full path from source data through every transformation to the fee applied and the invoice sent. "I can't go to a customer and say, here's a non-SOC 2 compliant billing solution," Michael said. "Every other billing solution on the planet is SOC 2 compliant." When compliance gaps appear, the cost runs beyond fines. There are regulatory filings that have to be amended, legal overhead, and time spent explaining and correcting. For Nasdaq's FMI customers, these same obligations apply. That means the platform Nasdaq delivers has to make defensible answers available on demand. Andrea described Databricks' role here as infrastructure-level auditability. Unity Catalog tracks every transformation, model, and dashboard: where the data came from, who touched it, when it changed, and what depends on it, automatically, without anyone needing to document it manually. "When an auditor or regulator asks," she said, "you have the answer." dbt adds the logic layer. Michael described pulling [a data lineage snapshot from dbt](https://www.getdbt.com/blog/what-is-data-lineage) to show a regulator the flow of a model from source to output. If there's a question about a specific calculation, the answer is in the documentation or the code itself. "dbt makes that part pretty effective and pretty easy," he said. It also helps customers building on top of Nasdaq's foundational models understand what those models mean and how they're calculated before making any modifications. ## Trusted data as the path to AI The AI conversation in financial services carries a particular weight. As Andrea framed it: "In financial markets, a confident but wrong answer is a risk event. It's not just an inconvenience." If a generative AI model hallucinates a product recommendation in a consumer app, a user gets a bad suggestion. If it hallucinates a position, a margin requirement, or a regulatory classification at a clearinghouse, that's a compliance breach. This is the gap the semantic layer is built to close. Michael described the root cause: AI gets things wrong not because models are incapable but because they lack context about the data they're working with. [The dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), which Michael called Nasdaq's "context layer," provides structured meaning alongside the data. When an agent accesses a metric, it knows what the metric means and how to apply it because the definition is explicit and governed, not inferred. dbt Labs [recently published a study on text-to-SQL accuracy](https://docs.getdbt.com/blog/semantic-layer-vs-text-to-sql-2026), comparing results with and without a semantic layer in place. The difference was stark: close to or at 100% accuracy when using semantic layer with both ChatGPT and Claude. Without it, agents querying raw tables produce confident answers that can be wrong in ways that are difficult to detect. Nasdaq is now building what Michael called an "agentic surface layer," investing in [Model Context Protocol (MCP)](https://www.getdbt.com/blog/mcp) tooling and skills that let customers build their own agentic workflows on top of the intelligence platform. The longer-term vision is a marketplace where Nasdaq, its partners, and its customers can share those workflows in a governed, verifiable way. "How do you do everything in a controlled environment such that you always know the result is guaranteed?" Michael asked. The foundation they've built is designed to answer that question, whatever the use case. When Jamie works through a financial services business case, the four themes he tracks—scale, productivity, governance, AI readiness—keep pointing back to the same thing: the work that satisfies a regulator today is the same work that makes AI trustworthy tomorrow. Lineage, contracts, semantic definitions, etc., aren't AI features. They're data infrastructure decisions. But they're the infrastructure decisions that determine whether AI agents can be trusted with the calculations financial markets depend on. Michael put it directly near the end of our conversation: "Delaying these initiatives is just becoming a bigger issue. There's a lot of focus on just trying to get there as quickly as possible." For Nasdaq's FMI customers, the option now exists to skip the three-year build and get there in six months. The combination of dbt and Databricks underneath the intelligence platform is a significant part of why that acceleration is real. [Watch the full recording](https://www.getdbt.com/resources/webinars/how-nasdaq-productized-a-governed-intelligence-layer-with-dbt-databricks-for-financial-market) to hear more of what Michael and Andrea covered. And if you're working toward the same kind of trusted data foundation, [talk to our team](https://www.getdbt.com/contact). --- --- title: "Ship smarter agents in production with dbt Agent Skills" description: "Just like humans, autonomous agents need faster feedback loops. Here’s how to develop them." url: "https://www.getdbt.com/blog/ship-smarter-agents-in-production-with-dbt-agent-skills" date: "2026-05-18" authors: ["Daniel Poppy"] categories: ["Insights"] --- # Ship smarter agents in production with dbt Agent Skills Coding agents are doing a tremendous amount of useful work today. Since Claude Code dropped last year, followed by Opus 4.5 and GPT 5.2, software engineering has very clearly passed a phase change. We've gone from copilot-style autocomplete to agents that can run end-to-end across the SDLC. But anyone who's tried to point one of these agents at a dbt project has hit the same wall I have. Ask a coding agent to build a new dbt model, and it'll happily make five or six changes across your DAG, then try to run the new model at the end. It breaks. The agent didn't know which columns existed, didn't iteratively run queries as it walked the DAG, and didn't think about lineage or contracts. The technology is marvelous. The agents simply haven't been taught how to do data work yet. Without a governed foundation—your actual models, lineage, contracts, and metrics—a coding agent is working from guesswork. It can write SQL that looks right and still return numbers no one can verify. That's what we're fixing with [**dbt agent skills**](https://docs.getdbt.com/blog/dbt-agent-skills). ## Coding agents are generalist agents, including for data We call them coding agents, but that framing undersells what they are. dbt agent skills can do all types of work. One of those types is data work, which is where most of you reading this probably want to put them. The catch is that the agents have been specialized for coding workflows. There's a long list of small tweaks and improvements that make them slot neatly into a software engineering loop. Data has its own additional bits that haven't been baked in by default: understanding [data lineage](https://www.getdbt.com/blog/what-is-data-lineage), respecting contracts, iteratively running queries to validate as you go, and knowing when to materialize what. A lot of teams have been layering those in by hand with AGENTS.md files. But there's a ceiling to how big an AGENTS.md can get before it becomes its own problem. ## What agent skills are, and what dbt's are doing Agent skills are [a protocol Anthropic released late last year](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) and donated to an open foundation. They're how you give an agent context into specific processes and workflows it needs to know about, packaged as markdown plus optional supporting scripts that the agent loads when relevant. We've taken everything dbt Labs has learned about analytics engineering and ported it into a series of agent skills, [in an open repo](https://github.com/dbt-labs/dbt-agent-skills) that works with any agent supporting the protocol. You can think of agent skills as **dbt best practices ported directly into your agent**. That includes: - Building a model iteratively rather than one-shotting it - Writing unit tests - Understanding your Directed Acyclic Graph (DAG) - Building your semantic layer - Debugging incremental models (which any longtime dbt user has spent their share of time on) The goal is to have the entire [Analytics Development Lifecycle (ADLC)](https://www.getdbt.com/resources/the-analytics-development-lifecycle) captured within skills, so the combination of a generalist coding agent and the dbt skills gives you a powerful data agent out of the box. That's part of the way there. The rest is your custom context, the things only your organization knows. Naming conventions, materialization choices, source quirks, gotchas. Skills are designed for that too: anyone can write one for their own project, and the strongest setups combine the general dbt skills with org-specific custom skills layered on top. ## How Factory scaled up with dbt Nikhil Harithas from [Factory](https://factory.ai/) has been building this out from scratch. Factory is in the business of bringing autonomy to software engineering through [Droids](https://docs.factory.ai/cli/configuration/custom-droids), their generalist agents that work across the SDLC, including the terminal, web, CLI, GitHub, and Teams. Nikhil's a field engineer there, and over the last few months, he's been standing up Factory's data posture using dbt as the central piece. Factory's materialization strategy changed as the team scaled up. Factory is a younger company, so the default was to materialize as late as possible, with a blanket rule that worked fine in the early days. As certain queries took long enough that things started getting expensive, and as the team wanted fresher data, that rule needed to change. The math, Nikhil notes, is both money and time, and also how many people you have to maintain the pipeline. Up until a few weeks ago, Factory was rebuilding entire tables on every single run because it was fine. The shift to incremental builds came because that was the only way to run more often without blowing up cost. Incremental builds are finickier than entire builds, so testing started to matter a lot more. As Nikhil puts it: "I wasn't as familiar with the [dbt test suite](https://docs.getdbt.com/docs/build/data-tests) until a couple of weeks ago, when I was like, okay, it's time to shore up all these kind of implicit contracts in these table definitions. We realized that every row has to have a distinct ID of some kind, depending on the table." Tests went on the most important tables first. The principle behind it: "How can you increase individual leverage as far as you can by systematizing as much as you can, by giving Droid the same kind of feedback you or I would?" ## Lessons learned from the build-out Nikhil’s top-line claim from the build-out: months of work that would have taken five or six people was done by one and a half people, in a couple of months. That alone is worth taking seriously. A few of the patterns he ran into: **Build vs. operate are different motions, and you need to teach the agent both.** Nikhil's framing: "What is it like to develop in dbt, and what is it like to operationalize in dbt?" Those things are related but slightly different. The historical reason agents have struggled with data teams more than they've struggled with generalist software engineering is, ultimately, a context problem on both fronts. Data is a mix of writing code and running operational workflows. It's closer to SRE work than pure software engineering in places. **Skill creep is real.** Early on, Nikhil ran into trouble with too many skills, which led to inconsistent triggering and ambiguity about which skill applied. The fix: Be intentional about how many skills you have, and make each one denser. Skills you're confident the agent will discover on its own can stay as habit-forming background. The high-value skills are the ones that anchor behavior on the most important tasks. **Say it louder in AGENTS.md.** Skills get auto-invoked sometimes, but you shouldn't bet on it. AGENTS.md is what's guaranteed to be in context, so Nikhil's pattern is to point at the relevant skills there explicitly: "You're going to be using [BigQuery](https://cloud.google.com/bigquery) and dbt and a few other vendors that [Extract, Transform, and Load (ETL)](https://www.getdbt.com/blog/extract-transform-load) data to us. You have to pay attention to the skills of these particular frameworks, because there's going to be how you do anything at all." **Build the repo so agents bump into skills naturally.** Even when a skill isn't auto-invoked, an agent grepping around to get its bearings should run into it. The mere mention of "dbt" anywhere in the project should surface the relevant SKILL.md in the search results. **Skill golf.** Doug Bady at dbt coined this. The practice: go through your skills and try to rip out everything you can to make them as tight as possible. Distractions in context cost you. **Documentation as a hook.** Factory now has a check that fires when someone changes a column or table. It requires documentation to land somewhere in the project before the change can merge. The agent doesn't have to write the docs by hand, but the structural requirement creates a flywheel where good behavior produces more good behavior. [To see these principles in action, watch a full demo of Factory’s Droids and dbt agent skills in our webinar.](https://www.getdbt.com/resources/webinars/ship-smarter-agents-building-for-production-with-dbt-agent-skills) ## Building faster feedback loops "Look at the last 20 years of software engineering,” Nikhil says.. “We’ve gotten faster because of feedback loops. It's easier to write tests, easier to write integration tests, logs are easier to look at. Humans are getting faster feedback loops. We have to give the same thing to agents." Nikhil's framing of where coding agents are today is worth leaning on. We're at a point, he says, where "if you can imagine it and if you are determined to build it, you can build it." That's a pretty magical thing to be able to say. But it feels like not everyone is experiencing that reality. The reason, in his view, is that while a lot of things that used to be difficult are now easy, some of the things that were hard are still hard. Integrations. Human context and assumed knowledge. Naming things, famously. If you're going to put one thing on your list this week, Nikhil's call was to figure out how to give your agents access to the most tools you can and the most amount of direct feedback you can. Nothing is more powerful than read-only access to the database. If an agent can't query the database, everything goes slower. The reframe he offered is the one I'd lead with: _What would need to be true for you to give an agent access to your database?_ That's the question. Maybe the work this week isn't a data engineering task at all. Maybe it's setting up your environment so you'd feel safe handing an agent that access. Once they have it, they fly. This is a singular moment in technological progression. The combination of generalist coding agents, dbt agent skills, the[ dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp), your own custom skills, and platforms like Droid is making it possible for very small teams to do work that used to require very large ones. Get involved. The [dbt agent skills repo](https://github.com/dbt-labs/dbt-agent-skills) is open, and we want to see what you build with it. --- --- title: "The dbt Developer Agent is now in Preview: the coding agent for analytics engineering" description: "dbt Developer Agent is now available in Preview—grounded in your dbt project so you ship faster without breaking downstream." url: "https://www.getdbt.com/blog/the-dbt-developer-agent-is-now-in-preview" date: "2026-05-06" authors: ["Chakshu Mehta", "Sam Ferguson"] categories: ["Product"] --- # The dbt Developer Agent is now in Preview: the coding agent for analytics engineering Every few years, the way we work with data evolves. We’ve gone from SQL editors to BI dashboards to conversational chat, and each shift moves data work forward. Now we’re entering a new phase defined by agents that can reason through tasks, plan, and act on their own. But despite being remarkably good at general software engineering, today’s coding agents struggle on dbt projects, because the project itself is the context and the guardrails. Without that grounding, they write SQL that looks right but references a column that doesn't exist, breaks a dialect rule, or silently breaks tests, contracts, and governed definitions three models later. The result is SQL that's syntactically correct and semantically wrong. Analytics engineers need an agent built for the job, grounded in your dbt project from the start. With today's launch, that's finally possible. The [**dbt Developer Agent**](https://docs.getdbt.com/docs/dbt-ai/developer-agent?version=2.0) is now available in Preview for dbt platform customers. It’s the next evolution of dbt Copilot, built directly in the Studio IDE, and it works across every file a change touches. That matters in analytics engineering, where the risk isn’t just a syntax error—it’s a change that slips past your project’s guardrails and breaks tests, contracts, or governed definitions downstream. No new tools to install. No context switching. Nothing to set up. It ships with [dbt Agent Skills](https://github.com/dbt-labs/dbt-agent-skills) and [dbt's product docs toolset](https://docs.getdbt.com/docs/dbt-ai/mcp-available-tools?version=2.0#product-docs) built in, so best practices and canonical answers are there while you build. Open dbt Studio to try it, or [read the docs](https://docs.getdbt.com/docs/dbt-ai/developer-agent?version=2.0#prerequisites) to learn more. ```json { "_key": "61029b8610f4", "_type": "heroVideo", "isModal": false, "url": "https://youtu.be/T5vRS9XSZSY" } ``` ## Agents that can think across your dbt project When we launched [dbt Copilot](https://www.getdbt.com/blog/introducing-dbt-copilot) in 2024, we were solving a real but bounded problem. Writing boilerplate, drafting tests, documentation, and YAML configs by hand slows everyone down without making anything better. dbt Copilot made those workflows much faster. But the thing that actually slows teams down isn't writing a model. It's what happens _around_ the model. You rename a column and something three dashboards deep breaks. You add a metric and spend an hour chasing down every semantic definition and exposure that needs to move with it. You open a PR and realize the last person who touched this file left your company six months ago. None of that is a coding problem. It's a data problem, and addressing that problem requires an agent that understands your data. During our beta, we spoke with dozens of customers, and kept hearing the same thing: they wanted an agent that actually knew their whole dbt project. One that could read the graph, understand the lineage, pick up on the contracts and the semantic definitions, and then make the change—without breaking the guardrails their teams depend on to keep data trustworthy. When something broke, they wanted it to troubleshoot the way a senior analytics engineer would: trace the failure, find the root cause, fix what's actually wrong instead of patching the symptom. So we built it. ## Meet the dbt Developer Agent The dbt Developer Agent lives right in dbt Studio, so you see every change in context directly in the IDE, right alongside the code you’re changing. In the IDE you can describe the change you want to make—like renaming a model, updating a column, or adding a metric—and the agent will analyze your full dbt graph. Not just file dependencies, but lineage, contracts, semantic definitions, tests, and governance. Then the agent drafts edits across those files and shows them to you as a sequence of reviewable diffs. You approve or reject each step, so only the changes you accept are saved to your project. > _“What really sets the dbt Developer Agent apart is its precision in identifying and resolving bottlenecks. It doesn't just suggest code; it understands our existing tests and lineage well enough to troubleshoot issues almost instantly. This has significantly reduced our build times and allowed us to scale our dbt project with total confidence in our data's trustworthiness. It’s the perfect balance of AI speed and strict architectural governance.”_ _Vishal Kaviraj, AVP Data Architect_ ## Grounded in your dbt project for faster, safer changes What separates the dbt Developer Agent from coding agents built for software engineering is two things: context and governance. Context is what makes an agent good at analytics engineering. The more of your project an agent can see, the safer its changes. Governance is what makes those changes trustworthy at scale, the tests, contracts, semantic definitions, and review controls that your team already depends on. Other agents built for data have pieces of that context. Warehouse-native agents read your schemas and write SQL within the warehouse they're built for. IDE-native agents (like Cursor, Claude Code) refactor across files and know your repo. Both are useful, and both are getting better every month. Some can even pull in pieces of your dbt project, but none of them are grounded in it. A dbt project isn't just SQL. It's years of accumulated decisions about how your organization's data should be shaped, tested, owned, and understood. That's the context that decides whether a change is actually safe, and that context lives in dbt. ## How it works The dbt Developer Agent runs on a loop, not a single shot. You describe what you want. It drafts the edits. It proactively asks to run `dbt compile` or `dbt build` to validate its own work. You stay in control with the ability to approve or deny each command as it goes. It sees the result, adjusts if it needs to, and keeps going until the change is ready for your review. it’s designed to help you move fast, but still work inside the same validation and review guardrails your data teams already trust. It also works with [Fusion](https://docs.getdbt.com/docs/fusion/about-fusion) out of the box. This ‌gives the agent a fast, local, deterministic feedback loop so it can validate changes (like missing columns, dialect rules, and downstream breakage) without waiting on expensive warehouse round-trips. A few things that make it great for developer workflows: - **It’s grounded in your whole project, not just the file you have open**. The agent understands your full dbt graph. When it writes or refactors a model, it understands what’s upstream, what’s downstream, and what a change means for the rest of the project. - **It keeps related files in sync.** A single prompt can produce coordinated changes across models, YAML configs, and documentation. If you rename a model, the refs follow. When you change a column, the downstream tests update with it. - **It ships with dbt Agent Skills.** Earlier this year, we launched [Agent Skills](https://docs.getdbt.com/blog/dbt-agent-skills?version=2.0), built by dbt Labs and the dbt community, that encode a decade of analytics engineering best practices. It brings the kind of knowledge that usually only lives in a senior engineer’s head, now available out-of-the-box to the agent, without any configuration. It also offers support for **project-level context skills**, enabling the Developer Agent to seamlessly detect and prioritize custom markdown skills alongside dbt Labs-managed ones. - **It ships with dbt's product docs.** So you can verify documented behavior and recommended patterns without leaving dbt Studio. - **It shows its work.** You see the agent’s reasoning and tool calls as it works, not just the final output, so if it takes a wrong turn you can see exactly where and why. You can copy the suggestion or open it directly in the editor. - **It keeps a human in the loop**. “Ask-for-approval” mode (the default) surfaces every edit as an inline diff before anything saves, while “edit-automatically” mode applies its work as it goes, with reasoning you can follow. You pick the right level of autonomy for the task, and nothing lands without your say. - **It validates as it goes.** A built-in comparison loop catches problems before you see the final diff, so agent-generated changes meet a higher bar than "the code compiles." - **It runs commands on your behalf.** Beyond just file edits, the agent can execute dbt commands, open pull requests, and handle the workflow steps around the change, not just the change itself. You approve each command before it executes, with options to allow it once for the session or deny it. ## On trust and autonomy We spent a lot of time thinking about how much autonomy a native dbt agent should have. The answer we landed on is that "it depends," and that the tool should support that flexibility rather than forcing a single workflow mode. - **“Ask-for-approval” mode (the default)** surfaces everything for review before saving. You see inline diffs, approve what looks right, and push back on what doesn't. Nothing happens without your say. - **“Edit-files-automatically” mode** writes and saves as it works, which can be right for ‌tasks where your team has enough confidence in both the agent and the change to let it run. Most teams will probably live somewhere between the two depending on the model, the stakes, and how familiar the work is. ## What you can do with the Developer Agent today The Developer Agent lives in the dbt Copilot panel in dbt Studio, so the whole loop of intent, change, validation, and review happens next to the code and lineage you're already working in. No context switching, no separate tool to learn. Here are a few of the things it handles well today: - **Refactor across files.** Describe the change, the agent reads your full graph, coordinates every file that needs to move, and surfaces the diffs for review. Rename a model and the refs follow. Restructure a mart and the downstream tests update - **Fusion migration.** For teams with projects blocked on conformance failures between dbt Core and Fusion, the agent classifies what's fixable, applies high-confidence fixes automatically, and surfaces what needs your input or is blocked at the engine level. - **Update models.** Describe a change in natural language and let the agent write or refactor the SQL, tests, and docs together, not separately. - **Enhance your semantic layer.** Add or modify metrics and dimensions with full project context. The agent knows your existing definitions and builds consistently with them. - **Migrate stored procedures.** Describe the logic you're moving off of and the agent translates it into dbt models, tests, and docs, with full awareness of where it fits in your existing graph. - **Create tests and docs.** Generate test coverage and documentation for existing models in a single pass, grounded in how each model is actually used downstream. - **Work without breaking your flow.** The Developer Agent lives in the Copilot panel, so you can stay in whatever file you're already editing and let the agent handle the related changes in the background. No switching to a chat tab to kick off work, then switching back to review it. > _We went from about 60 conformance errors to 7, using the Fusion migration agent. That's the difference between too hard and actually doable._" - Michael Fridolfsson, Data Architect, Brighte ## What's next This is just the beginning, and here are a few things we're actively building toward: - **Built-in data previews and data diffs** to improve validation and help you review the outcomes of agent-generated changes before they hit production. If your team works primarily in Claude Code or VS Code, the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) brings the same structured project context to the agent you already use. ## Get started The Developer Agent is available in preview for dbt platform customers with dbt Copilot enabled. Simply open dbt Studio, find the dbt Copilot agent pane, and describe what you want to build or change. “Ask for approval” mode is a good place to start: review the diffs, approve what looks right, and let the agent handle the rest. Not on the dbt platform yet? [Talk to our team](https://www.getdbt.com/contact) to learn more. --- --- title: "5 dbt MCP server patterns that work in production" description: "Five dbt MCP server patterns from real production use, including one that doesn't work the way you'd expect." url: "https://www.getdbt.com/blog/5-dbt-mcp-server-patterns-that-work-in-production" date: "2026-05-01" authors: ["Daniel Poppy"] categories: ["Pulse"] --- # 5 dbt MCP server patterns that work in production The Model Context Protocol (MCP) crossed 97 million monthly SDK downloads last December, the same month Anthropic handed it to the Linux Foundation. [The dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) is one of the more widely adopted data MCPs in that ecosystem. Here's what practitioners are doing with it. Five patterns. Two that work well, one that doesn't do what people expect, and two for when you want to go further. **1. Conversational analytics against governed metrics (works well)** Point a Claude or ChatGPT interface at the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer) through the MCP server. The server exposes your MetricFlow metrics, and the model queries them in plain language. When the model is grounded in governed definitions instead of writing raw SQL against undecorated tables, accuracy improves. **Where it shines:** metric-level questions against well-defined models. W**here to be careful:** complex multi-join queries still want a human in the loop. **2. First-draft documentation (works well)** Use the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) to generate documentation from column-level lineage and test coverage. The model reads the lineage graph, drafts model, and column descriptions, and you review them in the PR. Teams running this go from almost no coverage to most of their models documented. Setup takes a few hours. Treat the output as a starting point, not the final word. The review step is the whole game. **3. Real-time queries against very large tables (not what you'd expect)** Here's the one that bites people. Pointing an agent at full-scan queries against large, unpartitioned tables runs up warehouse costs and timeouts that you won't see coming at setup. The fix is simple: pre-materialize the semantic layer views you need, or add query guards before the agent runs anything. If your warehouse bill jumped and you're not sure why, start here. **4. CI/CD integration (advanced)** Wire the [dbt MCP server](https://docs.getdbt.com/docs/dbt-ai/about-mcp?version=2.0&name=Fusion) into your PR review. A pull request (PR) opens, the agent reads the diff, runs the impacted downstream tests through MCP, and writes a review comment. Teams that have set this up report much faster review cycles. This needs GitHub Actions or equivalent, plus a clear model impact map. There's a learning curve. It pays off. **5. Cross-tool orchestration (advanced)** Let an orchestration agent decide which dbt models to run based on upstream data quality signals. Source freshness fails, the agent skips the dependent models and pages the on-call engineer instead of running on stale data. Works in n8n and LangGraph. Not trivial to build, but you get a pipeline that degrades gracefully instead of quietly shipping bad numbers downstream. [**Check out dbt Wizard, the AI agent built for the way analytics engineers work.**](https://www.getdbt.com/product/dbt-wizard) --- --- title: "How Obie cut compute costs by 30%, reclaimed engineering hours, and built stronger governance" description: "See how Obie used the dbt Fusion engine and state-aware orchestration to cut costs, speed up pipelines, and scale with confidence." url: "https://www.getdbt.com/blog/how-obie-cut-compute-costs-by-30-percent" date: "2026-04-24" authors: ["Elaine Green", "Aika Zikibayeva"] categories: ["Product"] --- # How Obie cut compute costs by 30%, reclaimed engineering hours, and built stronger governance **** At Obie, an embedded insurance platform serving real estate investors, the data reality looked a lot like a typical fast-growing startup. With a lean footprint of just 130 employees and a recent, high-stakes acquisition, the pressure on the company's data platform was immense. Behind the scenes, infrastructure costs were climbing and starting to create friction across the business. The team adopted the dbt Fusion engine as their transformation engine and deployed state-aware orchestration (SAO) to remove operational drag and ensure confidence in the numbers the business depended on. When the team started, Obie's data layer was built on a Data Vault 2.0 methodology. This is a pattern designed for large enterprises. For a smaller fast-moving insurance tech company, it slows development and results in overhead that doesn’t match a smaller tech company’s size or pace. Obie made the call to re-architect entirely on Fusion, converting legacy models while preserving the underlying business logic. That migration is now over 90% complete. ## SAO reduces warehouse costs by ~30% and makes more frequent refreshes possible Previously, large portions of the pipeline were rebuilt whether upstream data had changed or not. As data volumes increased, that approach became expensive. Moving to [SAO on Fusion](https://www.getdbt.com/blog/announcing-state-aware-orchestration) meant running only what was necessary. Compute usage dropped, and warehouse spend became more predictable. The team could refresh data more frequently, from daily to every two hours, [without worrying about runaway costs.](https://www.getdbt.com/blog/dbt-compute-cost-reduction-fusion-state-aware-orchestration?utm_source=chatgpt.com) "We're saving at least 30% on compute costs, just from reusing models with state-aware orchestration,” said Tyson Doberneck, senior data engineer. Matt Karan, senior data engineer, added, “Knowing we're being proactive about costs gives leadership more confidence that we can have our data volume grow and still operate in a lean, young-company environment.” ## Fusion's orchestration and CI workflows recover up to 5 engineering hours per week With Fusion's built-in orchestration, version-controlled models in GitHub, and CI-driven staging environments that let engineers compare production vs. staging data before merging, pipeline interruptions decreased. Engineering time was freed to focus on analytics and product-facing work. ## Consistent metric definitions ensure consistent data The data team implemented best practices of data transformation: consolidating models, adding testing, and documenting definitions all within their Fusion project, the business can move faster because of consistent data. "Everything about Fusion has sped up my workflows,” said Karan. “I feel like it's just going to keep going in that direction. Eventually, I'll never have to leave my coding environment and be able to work with all the data in one pane of glass." ## What's next: Building the "Middleware" for agentic workflows and self-service analytics with dbt Semantic Layer Data development now feels noticeably smoother, and engineers can now focus on delivering business value. Next on the roadmap: an internal Slack bot that queries dbt's Semantic Layer to answer business questions on the fly. Because the bot references governed metric definitions rather than querying the database directly, the risk of hallucination drops significantly. The team is also evaluating dbt Mesh and the dbt MCP server as part of a broader push toward self-service analytics across Obie's 130-person organization. Doberneck concluded, “The dbt Fusion engine is a non-negotiable for me. With anything else in our stack, we could make a change. I would never switch out dbt.” --- --- title: "Using dbt with Databricks: Architecture decisions that determine success" description: "The cost of skipping dbt on Databricks compounds quietly. A solution architect explains what to watch for and when to act." url: "https://www.getdbt.com/blog/using-dbt-with-databricks-architecture-decisions-that-determine-success" date: "2026-04-22" authors: ["Keith Ludeman"] categories: ["Insights"] --- # Using dbt with Databricks: Architecture decisions that determine success _This guest post comes from Keith Ludeman, a solution architect at [Analytics8](https://www.analytics8.com/)._ Databricks gives you the platform. dbt gives you the structure, speed, and operating layer required to make it work for your whole organization. But getting the combination right requires decisions most teams underestimate, and the cost of getting them wrong compounds over time. I work with organizations at every stage of their [Databricks journey](https://www.analytics8.com/technologies/databricks-partners/)—teams standing it up for the first time, and teams that have been running on it for years and are hitting walls they didn't expect. The question I hear most often: do we really need dbt? The answer, in almost every case, is yes. But when you introduce it, and how, depends on where you are. Below I’ll walk you through: - The cost of going Databricks-native without a transformation framework - What dbt adds to a Databricks environment - Starting fresh: Why dbt belongs early in a Databricks implementation - Signals it's time to introduce dbt in an existing Databricks environment - Common questions and objections About dbt with Databricks - Practical lessons from real implementations ## The cost of going Databricks-native without a transformation framework Databricks gives you a lot of power right out of the gate. You can land data, shape it, orchestrate pipelines, run models, and serve analytics, all in one place. It is flexible enough to support almost any approach, which is exactly why teams lean into building things their own way early on. That usually starts small. An engineer writes a few transformations in notebooks. Someone wires up orchestration to keep things moving. As more use cases come in, more logic gets added, often by different people, each solving for the problem in front of them. Nothing feels wrong in the moment. The system is working. The problems show up later. As the environment grows, it becomes harder to answer basic questions. What does this model do? Where is this logic defined? If I change this, what breaks? There is no single way of doing things, so every answer depends on who built it and when. What used to feel flexible now feels unpredictable. That is where the friction starts to build: - **You lose consistency across the environment**. Different engineers structure transformations in different ways. Over time, it becomes difficult to follow the logic, compare approaches, or confidently reuse anything. - **Technical debt builds quietly**. What began as a handful of workflows turns into a network of dependencies that no one fully documents or owns. Testing becomes an afterthought; teams end up with basic row counts to confirm something ran, rather than automated, granular validation. Even small changes require extra caution, and new engineers need time just to understand how things fit together. - **Your team spends time maintaining the system instead of improving it**. Instead of focusing on analytics, you end up managing orchestration, testing patterns, and documentation standards on your own. The platform starts to demand attention rather than enable progress. - **Issues surface late, and often without clear signals**. Without consistent testing, lineage, and dependency tracking, problems do not show up where they start. They show up downstream, in dashboards, reports, or decisions, where they are harder to trace and fix. Most teams do not decide to build their own transformation framework. It happens gradually, through a series of reasonable choices. By the time it becomes a problem, you are not just dealing with messy code. You are dealing with slower delivery, rising costs, and a growing lack of confidence in the data. ## What dbt adds to a Databricks environment Before getting into what dbt adds, it’s worth clearing up a common misconception. dbt does not replace Databricks, and it is not another engine doing the work somewhere else. Your data still lives in Databricks, and your transformations still run there. dbt is simply the layer that defines how those transformations are written, tested, and maintained over time. That distinction matters, because most of the issues that show up in a Databricks-native environment are not about the platform itself. They come from how transformation logic evolves as more people start contributing to it. Without structure, that logic tends to live wherever it was first created. A notebook here, a pipeline there, maybe a few different patterns depending on who built what. It works, but it does not give the team a consistent way to understand or extend what is already in place. dbt changes that by giving you a single, shared way to define transformations: - You can follow how data moves from one model to the next. - You can see dependencies instead of guessing what might break - You have testing, documentation, and version control built into how the work gets done The shift is subtle at first, but it changes how teams operate. Engineers spend less time figuring out what already exists and more time building on top of it. Analysts are not blocked by how the data was originally created. And when something looks wrong, they can trace the logic themselves, propose a fix, and have it reviewed before anything changes in production. New team members do not have to reverse-engineer the environment just to contribute. The work becomes easier to reason about, which makes it easier to scale across teams. That same structure becomes critical when you layer [AI or advanced analytics](https://www.analytics8.com/services/ai-data-analytics-consulting/) on top of your data. Weak points show up quickly, and if transformations are hard to trace or validate, those issues do not stay contained. They surface in downstream outputs, where they are harder to catch and more expensive to fix. A structured transformation layer gives AI systems the context they need to return trustworthy answers and makes it easier to identify sensitive fields before they surface somewhere they shouldn't. ## Where the dbt platform advances this further As teams grow, the challenge usually shifts from how transformations are written to how they are run. In many dbt Core setups, orchestration becomes something you have to solve alongside everything else. Jobs are scheduled outside of dbt, dependencies are managed across tools, and over time you end up with another layer of logic that someone on the team is responsible for maintaining. The dbt platform changes that dynamic by pulling orchestration into the same place where transformations are defined. More importantly, it adds awareness of the data itself. If nothing upstream has changed, there is no reason to run the same transformations again. Instead of executing full pipelines on a fixed schedule, you run only what needs to run based on the state of the data. At smaller scale, that might feel like an optimization. At larger scale, it directly affects both cost and reliability. It also changes the day-to-day experience in ways that are harder to quantify but easy to feel over time. When you can trace dependencies more precisely, make changes across multiple models with confidence, and catch issues before anything runs in Databricks, the work becomes less about managing risk and more about moving forward. ## Why dbt belongs early on a Databricks implementation If your organization is implementing Databricks for the first time, the instinct is to start simple. Get Databricks running, prove it out, and add tooling later. That instinct makes sense. It is also where the retrofit tax begins. Most engineering teams have deeper SQL expertise than Spark or Python expertise. dbt is SQL-native, which means your team can start contributing immediately without a steep learning curve or relying on a small group of Spark specialists. It also changes your hiring options. SQL talent is easier and more affordable to find, and that advantage grows as your team scales. Most organizations are not starting from scratch. They are migrating from something else, whether that is SQL Server and SSIS, Google Dataform, or another SQL-based transformation environment. In those cases, the transformation logic is already written in SQL, and it ports cleanly into dbt. We have agentic migration accelerators that automate 70–80% of that work, compressing what would otherwise take years into weeks. That advantage depends on staying in a SQL-native model. If you go Databricks-native, you are rewriting that logic in a different paradigm. If you are building the platform now, this is the least expensive moment to make the right architectural decisions. A few foundational choices made at the start will determine whether your platform is easy to scale or expensive to untangle: - **Adopt Unity Catalog from day one.** Don’t take shortcuts. It is far easier to implement correctly upfront than to retrofit after you have already built on top of a structure that does not support it. - **Separate compute by workload from the start.** Segregating compute for ingestion, transformation, and reporting gives you cost visibility and performance control. You can see where costs are coming from and tune each workload independently. Without that separation, costs become harder to interpret and performance becomes harder to improve. - **Don’t reinvent orchestration with dbt Core.** Starting with the free version and building orchestration, testing, and observability yourself shifts the cost into engineering time. In practice, that effort often exceeds the cost of using a managed solution. The decisions that are cheap to get right at the start are expensive to fix later. Starting with dbt isn’t adding complexity. It’s avoiding it. ## Signals it's time to introduce dbt in an existing Databricks environment Some organizations bring dbt in from day one. Others start Databricks-native, run a successful small team, production environment, or proof of concept, and then hit a wall. These are the signals that tell you it is time to introduce, or reinforce, structure. - **Your team has grown past two or three people.** A small team can move quickly in Databricks without much formal structure. That breaks down once the team grows or when business analysts start contributing as analytics engineers. At that point, you are dealing with concurrency and change control. Multiple people working in the same environment without proper CI/CD and version control leads to conflicts, overwritten work, and compounding risk. dbt provides the structure that makes that collaboration manageable. - **You’ve added a second data domain.** A proof of concept built on sales data works. Leadership pushes to bring in supply chain. The moment you move beyond a single domain, both your data structures and your team structures need to evolve. Without that shift, each new use case becomes harder to add, maintaining what already exists gets more expensive, and both technical and organizational debt start to build. It is far easier to address this before the second domain is fully in production. - **The data team is still a bottleneck.** If business users are still waiting on IT for analytics, or if analysts are exporting data and rebuilding logic on their own instead of working from shared, governed definitions, your operating model has not kept pace with your platform. “I don’t want to be the chokepoint. I want to activate the business.” That is what you hear from data leaders at this stage. dbt makes it possible to bring business analysts into the analytics engineering process. Business analysts already know SQL. Very few know Spark. A SQL-based transformation layer helps close the gap between IT and the business, both technically and organizationally. When teams make the switch, the first thing they feel is speed—not in terms of compute, but in how quickly someone can answer a question. When an analyst asks why a number looks off, the engineer is not launching a research project. The logic is traceable from the final output all the way back to the source, model by model, in a way that is easy to follow and fast to navigate. Engineers also feel the difference in how safely they can make changes. With CI/CD guardrails baked into how dbt is structured, a change gets tested and reviewed before it touches production. That security changes how the team works—less caution, more momentum. ## Common questions and objections about dbt with Databricks ### **Can't we just do this natively in Databricks?** You can. Databricks is flexible enough to support almost any pattern. The tradeoff shows up over time in development effort, maintenance, velocity, and risk. dbt establishes the many patterns that teams often try (but fail) to build for themselves. Testing, documentation, version control, and standardized transformations are not new problems. dbt gives you a consistent way to handle them without reinventing that layer inside notebooks and pipelines. ### **When does dbt become necessary?** It rarely comes down to a single moment. What changes is the cost of continuing without it. The longer a team operates without structure, the more expensive it becomes to introduce it later. Patterns that are easy to establish early require rework once pipelines, dependencies, and team processes are already in place. If you are seeing the signals from the previous section, the cost has already started to compound. And if your organization is starting to think seriously about AI, that is another clear signal. The foundation of clean data, traceable logic, documented definitions that AI requires, is the same foundation that dbt helps you build. ### **Isn't dbt expensive?** dbt Core is free and can even be run within a Databricks job. That makes it a practical starting point. The real question is where the surrounding work lives. Orchestration, testing, observability, and maintenance do not go away with Core. They move to your team. The dbt platform shifts responsibility off your engineers. In many cases, the total cost balances out, while the operational burden drops. ### **How do teams decide between dbt Core and the dbt platform?** Most teams start with dbt Core because it is free and easy to spin up. That is a reasonable starting point, and for smaller teams it may be all you need. The conversation shifts, however, when complexity grows. When orchestration requires a separate tool, when you need more sophisticated dependency management, when the surrounding work starts pulling your engineers away from the analytics itself. At that point, the dbt platform tends to make sense. The licensing cost is relatively modest, and what you get in return (orchestration, testing, observability, all in one place), typically offsets what you were spending in engineering time to maintain everything separately. ## Practical lessons from real implementations Based on our experience implementing dbt with Databricks, here are a few practical recommendations to consider: - **Use each tool for what it does best.** Databricks is the right place to land and organize raw data. dbt is the right place to transform it into something the business can use. Teams that blur that line end up with an environment that is harder to maintain and harder to explain. - **Keep all transformation logic in dbt.** This is the single most common pattern we see in struggling implementations: logic split between notebooks and dbt models. It starts with one exception and compounds from there. Once logic lives in multiple places, lineage breaks down, debugging becomes a research project, and onboarding new team members takes significantly longer. - **Don't let logic creep into orchestration.** Orchestration should trigger work, not define it. When business logic gets embedded in orchestration tools, it becomes invisible to the rest of the pipeline. Changes to source systems break things in ways no one anticipates, because the logic affecting the data isn't where anyone thinks to look. - **Adopt Unity Catalog from the start.** Teams that skip it consistently pay for it later. It is one of those decisions that is low cost to get right early and high cost to retrofit once an environment has been built on top of something else. When these things are in place, the team stops spending its time maintaining pipelines and debugging failures. It starts spending that time on what matters—expanding capabilities, supporting new use cases, and engaging with the AI and advanced analytics work the business is asking for. _[Keith Ludeman](https://www.linkedin.com/in/keithludeman/) is a Solution Architect at [Analytics8](https://www.analytics8.com/) (a [dbt Labs Visionary Consulting & Services Partner](https://www.analytics8.com/technologies/dbt-partners/)) who specializes in designing and recovering complex data platforms across Databricks, Snowflake, and dbt, from migrations to governed transformation layers. He focuses on uncovering the root causes behind delivery challenges and putting the right structure in place so teams can scale without rework or technical debt. Known for stepping into high-risk situations and bringing clarity, he approaches each engagement as a long-term partner rather than a tool-driven implementer._ --- --- title: "dbt Labs Wins a 2026 Google Cloud Partner of the Year Award" description: "dbt Labs recognized for empowering thousands of Google BigQuery users to deliver trusted analytics and AI at scale" url: "https://www.getdbt.com/blog/dbt-labs-wins-2026-google-cloud-partner-of-the-year-award" date: "2026-04-21" authors: ["Elaine Green"] categories: ["Press"] --- # dbt Labs Wins a 2026 Google Cloud Partner of the Year Award **PHILADELPHIA – April 21, 2026: **[dbt Labs](https://www.getdbt.com/), the leader in standards for AI-ready structured data, announced today that it has received the 2026 Google Cloud Partner of the Year award for Data and Analytics: Data Pipelines and Governance. dbt Labs works together with Google Cloud to provide the foundation for an organization's transition to AI leadership and innovation. The combination of rich data warehousing capabilities and the democratization of complex data transformation removes technical barriers, enabling analysts and business leaders to accelerate their time-to-value. dbt Labs is being recognized for its achievements in the Google Cloud ecosystem, helping joint customers manage data at scale on Google Cloud and turn it into trusted, actionable insights with speed and efficiency. Thousands of organizations run dbt on Google BigQuery globally, an integration designed to accelerate the delivery of trusted analytics and AI. By consolidating data transformation into a single, unified tool, joint customers quickly gain increased operational efficiency through advanced orchestration features. dbt Labs empowers customers to manage and trust results, ensuring high-quality data is ready to power analytics and AI initiatives both today and in the future. “Every AI strategy needs to be underpinned by a standardized foundation and process to control, govern and document progress for high-quality, trusted results,” said Shawn Toldo, Vice President, Worldwide Partner Ecosystem at dbt Labs. “Together, dbt Labs and Google Cloud enable organizations to build that foundation for an AI-ready future. We are excited for the recognition and growing partnership with Google.” This recognition is the latest example of dbt Labs’ momentum since [launching](https://www.getdbt.com/blog/dbt-labs-launches-on-google-cloud-and-google-cloud-marketplace) on Google Cloud Marketplace one year ago. The partnership’s trajectory is driven by extensive global adoption and usage across diverse industries and a rapidly expanding community of active practitioners. Additionally, dbt Labs’ partner team earned two Google Partner All Star awards, reinforcing the deep collaboration and commitment to driving mutual success. “The Google Cloud Partner Awards honor the strategic innovation and measurable value our partners bring to customers,” said Kevin Ichhpurani, President, Global Partner Ecosystem and Channels, Google Cloud. “We are proud to name dbt Labs a 2026 Google Cloud Partner Award winner, celebrating their role in driving customer success over the last year.” By bringing Google AI capabilities into dbt workflows, joint customers gain the trustworthy, well-documented, governed foundation that reliable analytics and AI demand. To learn more about how dbt Labs and Google Cloud are enabling AI-ready data pipelines, watch the on-demand webinar “Building dbt Models Faster with Google AI” at [https://www.getdbt.com/confirmation/building-dbt-models-faster-with-google-ai-recording](https://www.getdbt.com/confirmation/building-dbt-models-faster-with-google-ai-recording). **About dbt Labs **Since 2016, dbt Labs has been on a mission to help data practitioners create and disseminate organizational knowledge. dbt is the standard for AI-ready structured data. Powered by the dbt Fusion engine, it unlocks the performance, context, and trust that organizations need to scale analytics in the era of AI. Globally, more than 80,000 data teams use dbt, including those at Siemens, Roche and Condé Nast. Learn more at getdbt.com, and follow dbt Labs on [LinkedIn](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=82576339&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fdbtlabs%2Fmycompany%2F&a=LinkedIn), [X](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=2299196361&u=https%3A%2F%2Fx.com%2Fdbt_labs&a=X), [Instagram](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=255235389&u=https%3A%2F%2Fwww.instagram.com%2Fdbt_labs%2F&a=Instagram), and [YouTube](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4561370-1&h=508887317&u=https%3A%2F%2Fwww.youtube.com%2Fc%2Fdbt-labs&a=YouTube). --- --- title: "Why metric definitions matter for reliable AI agents" description: "Learn how dbt's semantic foundation enables reliable, governed agentic analytics." url: "https://www.getdbt.com/blog/metric-definitions-ai-agents" date: "2026-04-21" authors: ["Joey Gault"] categories: ["Pulse"] --- # Why metric definitions matter for reliable AI agents ## The challenge of semantic ambiguity [dbt](https://www.getdbt.com/product/what-is-dbt) agents operate fundamentally differently than human analysts. When a human encounters ambiguity in a metric definition, they can apply context, ask clarifying questions, or make informed assumptions based on institutional knowledge. AI agents lack this intuitive understanding. Without precise definitions, they generate inconsistent results that undermine trust and create operational risk. Consider a seemingly straightforward metric like "monthly revenue." Across different departments, this could mean revenue recognized in a given month, revenue booked that month, revenue from contracts starting that month, or revenue adjusted for returns and refunds. A human analyst working with the finance team understands which definition applies in context. An AI agent querying data autonomously does not. When multiple agents operate across different teams—a sales agent analyzing pipeline performance, a finance agent generating forecasts, and a customer success agent evaluating retention—semantic inconsistency creates a compounding problem. Each agent might calculate "monthly revenue" differently, producing conflicting outputs that require manual reconciliation. This defeats the purpose of autonomous analytics and erodes confidence in agent-generated insights. The scale of this challenge becomes apparent in modern data environments. Organizations now work across an average of 400 data sources, with nearly one in five enterprises managing more than 1,000 sources. In these complex ecosystems, the same business concept might be represented dozens of different ways across systems. Without a shared semantic layer that provides consistent definitions, agents amplify rather than resolve this fragmentation. ## Structured context as the foundation for agency For AI agents to operate safely and effectively in enterprise environments, they require more than instructions and advanced language models. They need structured context: the schemas, semantics, relationships, permissions, and lineage that describe how data works within an organization. Structured context equips agents with three essential capabilities. First, it provides memory through metadata, enabling agents to understand what data assets exist and how they relate to one another. Second, it establishes boundaries through clear definitions, permissions, and rules that prevent agents from operating outside established guardrails. Third, it enables useful actions by providing validated tools and interfaces for reading and writing data safely. Metric definitions sit at the heart of this structured context. When an agent needs to answer a question about customer churn, it must know precisely how "churn" is defined, which data sources contain the authoritative calculation, what business rules apply, and who has permission to access the underlying data. Without this semantic foundation, agents resort to guessing or hallucinating definitions, producing unreliable results. dbt excels at creating this structured foundation through [data transformation](https://www.getdbt.com/product/develop) workflows that convert raw data into analytics-ready models with built-in testing, lineage tracking, and semantic meaning. By defining metrics consistently within dbt models and exposing them through the [dbt Semantic Layer](https://www.getdbt.com/product/semantic-layer), organizations create a single source of truth that both humans and agents can reliably query. ## The cost of poor definitions The consequences of weak metric definitions become severe when agents operate autonomously at scale. Poor data quality is cited as the primary reason AI projects fail to deliver expected value, and organizations lose an average of [$12.8 million annually](https://www.ibm.com/thought-leadership/institute-business-value/en-us/report/data-quality) due to data quality issues, with some companies losing as much as 6% of annual revenue from flawed AI outputs. High-profile failures illustrate the risk. Airlines have faced legal action when chatbots promised refunds based on hallucinated policies. Lawyers have submitted briefs citing fictional cases generated by AI systems. News organizations have published AI-generated travel content directing readers to unsafe destinations. Even the most advanced language models hallucinate at significant rates when operating without proper grounding in structured, validated data. In the context of business analytics, these failures manifest as agents that confidently report incorrect metrics, make recommendations based on flawed calculations, or trigger automated actions using inconsistent business logic. When an agent autonomously adjusts pricing based on a miscalculated margin metric or sends customer communications based on an incorrect churn definition, the operational and reputational damage can be substantial. Regulatory frameworks are making these risks explicit. Under the EU AI Act, particularly Articles 10 and 27, organizations deploying high-risk AI systems must demonstrate that their data is complete, accurate, representative, and error-free. This includes comprehensive documentation of data sources, quality checks, and bias mitigation measures. Metric definitions are a core component of this compliance obligation: organizations must be able to prove that their AI systems are calculating business-critical metrics correctly and consistently. ## Governance through definition Effective governance for AI agents cannot be bolted on after deployment. It must be embedded in the [data transformation](https://www.getdbt.com/product/develop) layer where metrics are defined and calculated. This approach ensures that governance policies flow automatically through dependent models and that agents inherit the correct definitions and access controls. When metrics are defined centrally in dbt, changes propagate consistently across all downstream uses. If the definition of "active user" changes to reflect new product features, that update flows automatically to every dashboard, report, and agent that references the metric. This eliminates the drift that occurs when definitions are scattered across multiple systems or hardcoded into individual queries. Column-level security and row-level access controls become particularly important for agents. Unlike human users who might access a dashboard showing aggregated metrics, agents often query underlying data directly. A conversational analytics agent responding to a sales manager's question about team performance should only access data for that manager's region and team members. These access boundaries must be defined at the metric level and enforced consistently regardless of how the data is accessed. The [dbt MCP (Model Context Protocol) server](https://docs.getdbt.com/docs/dbt-ai/about-mcp) provides a standardized interface for exposing dbt models, lineage, and semantic context to AI systems while maintaining fine-grained, policy-aware access controls. This enables agents to discover available metrics, understand their definitions and lineage, and query them safely within established governance boundaries. ## Enabling multi-agent collaboration As organizations move beyond single-purpose agents to multi-agent architectures, consistent metric definitions become even more critical. In these systems, specialized agents handle specific functions and collaborate through orchestration layers to complete complex tasks. Consider a scenario where a discovery agent helps a business user identify relevant datasets, an analyst agent generates insights from those datasets, and a developer agent creates new data models based on the findings. For this workflow to function reliably, all three agents must share a common understanding of the metrics involved. If the discovery agent surfaces a "customer lifetime value" metric that the analyst agent calculates differently, the entire workflow breaks down. Event-driven architectures for multi-agent systems depend on semantic consistency. When one agent publishes an event indicating that a key metric has crossed a threshold, downstream agents must interpret that metric identically to respond appropriately. This requires metric definitions to be versioned, documented, and accessible through shared interfaces that all agents can query. Organizations implementing multi-agent systems should treat metric definitions as contracts between agents. Just as microservices rely on well-defined APIs, agents rely on well-defined metrics. Changes to metric definitions should follow the same rigorous change management processes as API changes, including versioning, deprecation notices, and backward compatibility considerations. ## Practical implementation for data engineering leaders Building a metric definition framework that supports reliable AI agents requires deliberate architectural choices. Data engineering leaders should focus on several key practices. Start by mapping all business-critical metrics and documenting their definitions comprehensively. This includes not just the calculation logic, but also the business context, data sources, refresh frequency, known limitations, and ownership. These definitions should live in version control alongside the dbt models that implement them, creating a single source of truth that evolves with the business. Implement comprehensive testing for metric calculations. [dbt's testing framework](https://docs.getdbt.com/docs/build/data-tests) enables data teams to validate that metrics are calculated correctly, that underlying data meets quality standards, and that changes don't introduce regressions. For AI agents, these tests serve as guardrails that prevent autonomous systems from operating on flawed data. Establish clear ownership and approval processes for metric changes. When an agent relies on a metric definition to make autonomous decisions, changes to that definition have operational implications. Metric owners should be identified, change requests should be reviewed by stakeholders, and impacts should be assessed before deployment. Expose metrics through a semantic layer that provides a consistent query interface for both humans and agents. The [dbt Semantic Layer](https://docs.getdbt.com/docs/use-dbt-semantic-layer/dbt-sl) enables organizations to define metrics once and query them consistently across tools, eliminating the proliferation of slightly different metric implementations that creates semantic drift. Monitor how agents use metrics in production. Observability for agentic systems should include tracking which metrics agents query, how they interpret results, and what actions they take based on those metrics. This visibility enables rapid intervention when agents misinterpret metrics and creates feedback loops for improving definitions. ## The path forward The shift toward agentic analytics is accelerating. According to the [IBM Institute for Business Value](https://www.ibm.com/thought-leadership/institute-business-value/en-us/report/agentai), 70% of executives consider agentic AI critical to their future, and 61% of CEOs report actively deploying or scaling AI agents. The agentic AI market, valued at $5 billion today, is projected to reach $50 billion by 2030. Organizations that establish rigorous metric definition practices now will be positioned to deploy autonomous analytics systems with confidence. Those that treat metric definitions as an afterthought will struggle with unreliable agents, inconsistent outputs, and erosion of trust in AI-generated insights. For data engineering leaders, the imperative is clear: invest in the semantic foundation that makes autonomous analytics possible. [dbt's semantic layer](https://www.getdbt.com/product/semantic-layer) provides the transformation framework for defining metrics consistently, testing them rigorously, and exposing them through interfaces that agents can query reliably. Explore [dbt's agent capabilities](https://docs.getdbt.com/blog/dbt-agent-skills) and the [dbt documentation](https://docs.getdbt.com/docs/introduction) to learn how to implement these practices at your organization. The organizations that will thrive in the era of agentic analytics are those that recognize metric definitions not as a documentation exercise but as critical infrastructure. By treating metrics as first-class data products with clear ownership, rigorous testing, and consistent governance, data engineering leaders create the foundation for AI agents that augment rather than undermine analytical capabilities. ## AI agent FAQs **What is the difference between AI agents and traditional business intelligence tools?** AI agents operate autonomously and lack the intuitive understanding that human analysts possess. When human analysts encounter ambiguity in metric definitions, they can apply context, ask clarifying questions, or make informed assumptions based on institutional knowledge. AI agents cannot do this: without precise definitions, they generate inconsistent results that undermine trust and create operational risk. Traditional BI tools are typically operated by humans who can interpret and contextualize data, while AI agents query and analyze data independently, making structured context and clear metric definitions essential for their reliable operation. **What analytics are available to assess AI agent performance against business objectives?** Monitoring how agents use metrics in production is essential for assessing their performance. Observability for agentic systems should include tracking which metrics agents query, how they interpret results, and what actions they take based on those metrics. This visibility enables rapid intervention when agents misinterpret metrics and creates feedback loops for improving definitions. Organizations should monitor agent outputs for consistency, validate that agents are calculating business-critical metrics correctly, and ensure that autonomous decisions align with established business logic and governance policies. **How do AI agents improve consistency and trust?** AI agents improve consistency and trust when they operate on a foundation of structured context with precise metric definitions. By defining metrics centrally in transformation workflows, changes propagate consistently across all downstream uses: every dashboard, report, and agent references the same definition. This eliminates drift that occurs when definitions are scattered across systems. When metrics are treated as contracts with clear ownership, rigorous testing, and consistent governance, agents can reliably query a single source of truth, producing consistent outputs across different teams and use cases rather than conflicting results that require manual reconciliation.